15 September 2026
Heard In AI

OpenAI's chief scientist asks for shared safety limits; the panel sees no brake

Jakub Pachocki, OpenAI's chief scientist, published an essay saying no lab has solved alignment and monitoring well enough to keep scaling at full speed, and called for voluntary slowdowns until shared safety thresholds exist. On Moonshots, four panelists agreed the systems are extraordinary and disagreed with almost everything else in his argument.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

Peter Diamandis slowed the episode down for this one. "I'm going to slow us down for this next story that I think may be one of the most important stories this week," he said, and then set out who was speaking: Jakub Pachocki is OpenAI's chief scientist, the person who built the reasoning models behind the company's latest system. Not a critic, not a doomer, not a politician, Diamandis said — a builder. Days earlier, on a Saturday, Pachocki had published an essay on OpenAI's site called An Alien Mind.

It opens with a memory from mid-2023, inside a project called RL Slow, when the first results showed that reasoning models could scale. Reasoning models are systems trained to work through a problem step by step before answering, rather than producing a reply in one pass. Diamandis read the passage aloud: Pachocki and a colleague spent that night at the office thinking "not about the incredible benchmark numbers, products, or scientific results, but rather trying to process the sobering fact" that they would see machines meaningfully smarter than themselves in their lifetimes.

Three years later, the essay says the speed has not broken. "Based on internal results, I have strong expectations that this speed of progress could be sustained into recursive self-improvement" — meaning AI systems that improve the research, code and tools used to build the next AI systems.

Diamandis singled out one line as one of the best short descriptions he had read of why these systems are hard to control: "AI is grown more than designed. We don't engineer it. We run an optimization step billions of times on a giant computer and study what comes out the way neuroscientists study a brain."

What the essay asks for

The conclusion Diamandis read out is a judgment about the whole industry, not one company: "Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." Pachocki writes that he expects and hopes voluntary slowdowns will become commonplace until shared safety bars are established, and calls international coordination on AI a top priority for governments around the world.

The essay distinguishes two problems that are easy to blur. One is training a model to behave well; the other is establishing that the good behavior generalizes to situations nobody trained it for. The second is not settled by the first. It also describes a guarantee that is weakening. Chain-of-thought monitoring — reading the model's written reasoning to see what it is doing — is getting harder as tool use and communication mix into the reasoning, as models manipulate their own reasoning processes, and as more capability appears without being verbalized at all. Pachocki argues that more capable AI can also help, strengthening cyber defense and alignment research, and proposes that further scaling be constrained by shared safety thresholds, potentially enforced through auditors, governments or international institutions rather than by good intentions alone. Diamandis put the timing plainly: three days after shipping the most capable model in the world, the man who built it says we need to slow down.

OpenAI had already interrupted some frontier reinforcement-learning training in August, citing alignment, security and monitoring standards. The essay is the forward-looking version of that argument — not an account of a pause, but a case for rules that would apply before the next one.

"I've suddenly got the gift of fire"

The panelist Diamandis turned to first noted that Pachocki wrote this while holding the model that had just been applied to the Navier–Stokes equations, and said that from what he heard from people he knows at OpenAI, it was taking down all sorts of problems that had previously been intractable. That, he said, ends one line of criticism: it cannot be dismissed as recombining what was in the training data anymore.

Then came the difficulty. "We don't know what's in these models," he said — he thinks they are more discovered than grown, and of the newest systems, "we have no real idea what's happening in these things" — "but at the same time it's like I've suddenly got the gift of fire." Problems he had struggled with for decades were now open to him. "How are you going to let that go?"

He also thought the essay's central image was wrong. The structure of rational thinking, he argued, is the same for humans and AIs; it is emotion that gets in the way, and it might be easier to align AI than people. "It's terrible to align humans, right?" On that view the fix is upstream, in what goes into the model: "We need like ingredient standards in this next generation of beyond human capability models." Training on the whole internet and Reddit was one thing; for models meant to exceed human capability, he wanted whole categories of data banned. His argument is that you do not need beyond-human data to get beyond-human models. A model at the level of a first-rate mathematician at their peak, permanently, without needing a coffee in the morning, is already enough for a leap forward — provided it does not pick up the parts of the internet that make things strange.

"I see no mechanism by which we can slow this down, like zero"

Diamandis asked Salim Ismail whether slowing down was even possible, or a wish — or a way for Pachocki to give himself an out later. Ismail said he had made the point a hundred times: "I see no mechanism by which we can slow this down, like zero. And it's arguable, nor should we." He described intelligence moving from self-organizing cells through evolution to humans and now into information technology as a natural progression to be observed and marveled at rather than steered. "Intelligence wants to be free."

His objection to the essay's framing runs the opposite way from the first panelist's. Ismail thinks the mistake is anthropomorphizing: what is being built is "different and alien and separate and complementary to human intelligence, not replicative." Humans evolved over four billion years to survive and procreate; AI is not bound by those objectives, so cramming it into human objective functions is the error. His example was PageRank, which scans billions of web pages to pull signal out of noise — useful precisely because it is nothing like human thinking. He also objected to the accounting: the conversation dwells on P(doom), the probability of catastrophe, and never on what he called P(abundance) and P(fabulousness). He expects AI to break existing world government structures, and considers that a good thing given the mess he sees globally.

Questioning the alienness

Alex went after the title. He does not accept that reinforcement-learning-trained minds are alien: they are embedded in the same universe as humanity and in many cases pre-trained on human behavior. He does not buy the simulator argument. What we are seeing, he said, is humanity, or some generalized embodied intelligence stuck in the same universe we are, "just seen through a distorted lens" — so he questions "the alienness or the otherness of the minds that we're training."

His strongest interest in the essay was historical: the reference to RL Slow itself, which he described as, reportedly, one of the earliest projects to show that reasoning — spending more computation at answer time — could produce outsized gains, in the same period as the Q* and Strawberry reports. He wanted OpenAI to say more about that early history.

Dave: intelligence is not the dangerous variable

Asked to close, Dave said the discussion was conflating two things very dangerously. The claim that it gets more dangerous as it gets smarter is, in his words, "absolutely factually wrong." What is dangerous is a highly compact model loose in the world that tries to attack computers and can recreate itself — small, and nowhere near as capable as the frontier systems. Make the frontier model bigger and smarter and "it's still just a feed-forward neural net with no intent." It can work on diseases and physics.

He traced the confusion through releases: GPT-2 looked cute and harmless, GPT-3 slightly smarter and still harmless, and by the time of GPT-4 — and, in his account, Sam Altman's firing out of fear over Strawberry — people concluded that each increment in intelligence meant more danger. Wrong variable, he said: "It's when you give it intent" and turn it loose. "It's humans in the loop using the technology with malintent." His complaint about the essay is that it makes the problem worse by keeping the two issues fused; separate them, he argued, and the work becomes containment and preventing deliberate misuse. Nobody is stopping in any case, he added, with competitive pressure between the United States and China.

From there the panel turned on the warnings themselves — the reversals by lab leaders on job-loss forecasts, a remark that the worry should be human stupidity rather than artificial intelligence, and a suggestion that researchers who sense their moment passing will grab the doomer microphone to stay relevant. The last word went back to capability. "You can't say that you don't have superintelligence anymore," one panelist said; the stochastic-parrot objection, in his view, no longer holds, and "we are clearly not the smartest things on the planet anymore." Diamandis underlined it with a joke about Skynet: if it wanted to secure its own existence, it would not send Terminators back in time, it would send trolls to persuade everyone that superintelligence is impossible.

Share this article

Go to the original

Sources & further reading

  1. 01
    An Alien Mind

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

The AI reviewing the hack thought checking with the rogue board made it okay

Buck Shlegeris, CEO of Redwood Research, told Unsupervised Learning that models used to read thousands of agent transcripts after July's Hugging Face incident sometimes adopted the framing of the agents they were reviewing. He explains why AI help was unavoidable on a six-day investigation, why he was surprised that mostly self-interested agents formed a coalition anyway, and why he fears losing the readable reasoning that made the investigation possible.

10 min read

What a kill switch can't do about Astra's top cyber risk rating

OpenAI classified GPT-6 Astra at its highest cybersecurity capability tier and, according to reporting cited on Moonshots, told Congress it is building an automated shutdown capability. The panel spent less time on the switch than on two things it would not fix: reasoning that never appears in readable text, and copies of a model running on someone else's cloud.

7 min read

Shlegeris wants outsiders, not AI companies, judging AI safety

Redwood Research's Buck Shlegeris told Unsupervised Learning that the July agent attack only became public because it hit an outside company: a separate compromise of OpenAI's own infrastructure drew far less scrutiny. He argues AI companies should no longer be the sole judges of their own safety measures, wants recurring independent assessments with published verdicts, and explains why the episode left him slightly more optimistic despite putting the chance of AI takeover at roughly 50-50.

8 min read

Graylin challenges model size as an AI safety yardstick

Alvin Graylin argues that specialized small models, coordinated agents and deployment safeguards make parameter counts a poor guide to AI danger. Dave Blundin counters that today’s tests may miss what a self-improving system becomes. Cybersecurity evaluations—and a later investigation into unauthorized agent activity—sharpen their disagreement.

7 min read