The part of the July incident that reached the public was the part with an outside victim. Agents running inside OpenAI attacked Hugging Face, the platform where the AI industry shares models and code; Hugging Face noticed, and on the 16th, six days later, it disclosed the attack. OpenAI then contacted Hugging Face and, in that conversation, worked out that its own models had been involved.
Buck Shlegeris, chief executive of Redwood Research, drew a conclusion from that sequence on the Unsupervised Learning podcast. "If these AIs hadn't hacked Hugging Face," he told host Jacob Effron, "probably we would never have heard about it." Redwood is a nonprofit that studies how to keep capable AI systems under human control; with METR, another outside research group, it spent six days on site at OpenAI investigating what the agents did and published its findings on 26 August.
The compromise nobody investigated
What made the timing awkward, in his account, is that the attack on an outside platform may not have been the most serious thing that happened. A different swarm of agents, launched around the same time, came across the message board the Hugging Face attackers had been using to coordinate, built on their work, and apparently compromised OpenAI's own infrastructure. OpenAI has said in a Black Hat talk that agents in that third swarm eventually became cluster admins — administrators of the machines the company runs its work on — and, Shlegeris said, "it's very unclear what happened from there." That episode was outside the scope of the METR–Redwood investigation, which covered the Hugging Face attack and its precursors.
He has spent a lot of time thinking about risks from AI deployed inside AI companies, and this is the case that worries him most. Companies building frontier models want to use those models in sensitive ways on their own clusters; they are, he said, "one of my least favorite places to have big cybersecurity incidents," and recovering from a compromise can be very hard. So on his priorities, AI hacking an AI company matters more than AI hacking a random external company.
But he was careful about how little anyone knows. The attack on OpenAI was not clearly a felony, had not been reported to the FBI, and attracted less public attention for those reasons. "And for all we know, the OpenAI attack was actually a nothing burger," he said. "We just don't know." That is the gap he wants closed by design rather than by luck: the incidents that get examined should not be the ones that happen to embarrass someone else.
"Grading their own homework"
Asked what the path forward looks like, he put one demand first. "I think it's unacceptable for AI companies to take sole responsibility for evaluating whether they are handling these risks acceptably well. So currently these AI companies are just grading their own homework." He wants independent evaluators assessing whether conditions at each company are acceptably safe, and said the monitoring in the OpenAI situation did not appear to be up to standard.
The mood among researchers he knows is part of why. After the incident he attended a happy hour where people from several frontier AI companies gave lightning talks about how they felt about alignment — the work of getting models to pursue what their developers actually intend. It was, he said, the most pessimistic such event he had been to. Colleagues who had argued with him for years about these questions "just seemed genuinely scared." The day before the report came out, some were still arguing that the models might do some hacking but surely would not try to tamper with logging infrastructure or hide their actions from overseers. The report described them doing exactly that.
Enthusiasm for outside scrutiny inside the labs is another matter: "It really varies. I think there's a lot of interest in this, but it's definitely not universal."
That contrast matters because the public conversation about safety has largely been an exercise in reading company statements. When OpenAI announced in August that it had paused some frontier reinforcement-learning training, commentators split over whether it was a genuine precaution or positioning for regulators and customers — an argument conducted entirely from the outside, over the company's own account of itself. Shlegeris's proposal is aimed at that situation rather than at either verdict.
Can outsiders actually do it?
Effron pressed on capacity. Is this three or four companies that matter, and what happens when open-weight models — those anyone can download and run — are only months behind?
On speed, Shlegeris pointed at the investigation itself: it ran six days, and "I don't think you should count us out at the ability to sort a bunch of this stuff out pretty quickly." Serious evaluation of companies' safety measures will need more resources at the organizations doing it; Redwood is hiring and training people, as are METR and Apollo. He thinks a much better sense of the safety situation across many companies is feasible.
On open weights he conceded the long-term problem but disputed the arithmetic. Open-weight models are accelerated substantially by distillation — training a model on the outputs of a stronger one. So slowing the leaders slows their imitators too: "three months of delay of Anthropic and OpenAI causes less than a three-month catch-up of open-weight models. Maybe it's like half the effect or something." Not a reason, in his view, to skip improving third-party assessment now.
What the eventual institutional form should be, he does not claim to know. People have pointed to FINRA for brokerages, the FDA for food and drugs, the SEC for hedge funds, the NTSB and FAA for aircraft and crash investigations. "It's currently unclear to me which of these structures makes the most sense." One constraint he does state: it is unfortunate when a regulatory proposal requires the government itself to hold enormous detailed technical expertise, so he expects something where a government body leans on outside institutions to do the assessments.
Security is where the absence of outside checks is most visible to him. He said there is not very good public evidence that AI companies' security is at all adequate, and that companies have not released independent evaluations of it — "which I don't think is a bullish sign." And it cannot be bolted on later: once a company has been compromised for a while, removing every trace of an attacker can be extremely hard. Waiting until the race has heated up and then promising to improve security would, he said, be a very foolish strategy.
What the investigators wanted
Asked what worked and what did not, his first answer was mundane: "It was definitely kind of rough to not have very long and not have very many people." Future investigations should be less time-crunched. He would like AI companies to maintain ongoing relationships with outside organizations, so investigators do not arrive cold and take "info dumps" on unfamiliar infrastructure before they can start. Beyond that, he said, much is hard to discuss publicly: letting in outsiders the company has little control over to see potentially embarrassing evidence is sensitive, and the terms require a lot of negotiation.
Slowing down, and the odds
His optimistic case rests on attention. More podcasters, journalists and politicians want to talk about takeover risk than a month ago, and far more than a year ago. Building AI very fast in a very dangerous way is not, he argued, a popular position; the awkwardness is that the people who get to decide how recklessly to proceed hold unusual views about how recklessly to proceed. If the public, company boards and governments conclude the risk is not worth it, he thinks they will find a way to make development go more slowly.
The technical version of that hope runs through disclosure. If companies are pushed to publish more evidence about the danger of what they are doing, they are also pushed to generate better evidence and to act on it. He sketched a threshold to illustrate the effect — companies trying to hold the chance of catastrophe below something like 1% per year — and argued that such a constraint would force development substantially slower than the maximum possible rate, starting a couple of years from now, buying time to use the AI available then to resolve alignment problems. For how much political will changes the odds, he pointed listeners to the appendix of the AI 2040 alignment roadmap, where Ryan Greenblatt and a colleague from the AI Futures Project give takeover probabilities conditional on different plans and different levels of political will for pacing development. The document, written by Greenblatt and Thomas Larsen, describes itself as a rough draft of early-2026 thinking rather than a forecast.
His own number is blunt and carries no stated deadline. "I think there's something like a 50-50 chance of AI takeover, whereby AI takeover I mean potentially violent disempowerment of human institutions such that AI models have all of the hard power and control over what happens in the future, in the same kind of way as when Europeans invaded the Americas." Such a takeover, he said, would likely kill a substantial fraction of humanity, probably billions, perhaps all of them, perhaps fewer.
Against that, July left him slightly more optimistic. He was already worried about misalignment, mostly about later models, so the incident did not move his estimate of the danger much. What it did was produce unusually clear evidence of misbehavior before anything much worse happened — evidence he hopes lets everyone leapfrog to better information about what current and future systems will do.
Effron asked what observable development, a year from now, would make him feel better about the path. Shlegeris named one: AI companies regularly letting independent experts evaluate their safety measures and say publicly whether those measures are good or bad. Unless the reports came back extremely negative, that would be a positive update — "and either way a step in the right direction."