Peter Diamandis said the tone of the post mattered enough to read the whole thing. OpenAI had just claimed a breakthrough on the Navier–Stokes Millennium Prize problem — one of a famous list of unsolved problems in mathematics, this one about whether the equations describing how fluids move always behave — using what the company described as roughly 10,000 AI agents working at once. Sam Altman's reaction, as Diamandis read it on Moonshots with Peter Diamandis: "The world has extremely capable models now. I did not expect a result of this magnitude to happen so soon. We've been talking a lot about the need to pace progress to ensure safety. For me, this is the strongest evidence yet for that urgency."
Diamandis put the obvious question to the table. Is this genuine alarm? Is it marketing? What is it?
Nobody on the panel wanted to stop. Almost everybody wanted something from the labs, and they disagreed sharply about what.
Three signals in five days
Diamandis laid out the sequence. On the Saturday, OpenAI's chief scientist Jakub Pachocki had published An Alien Mind, arguing that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer, and calling for voluntary slowdowns and shared safety thresholds. Alignment, in this conversation, means getting a system to pursue what its makers and users actually want, including in situations nobody trained it for. Then came Altman's post. Then, on screen, a resignation notice from a researcher Diamandis named as Jacob Coxon, who he said had spent three years on pre-training at both OpenAI and Anthropic. The charge, quoted: "neither company is acting responsibly in the race towards self-improving superintelligence." The labs were pursuing it, the researcher wrote, while internally believing sufficiently powerful systems could pose catastrophic risks. He called the competition "gambling with our lives."
"Three signals from inside the labs in five days," Diamandis said.
A fourth followed: a reply from Evan Hubinger, Anthropic's alignment science lead, which Diamandis also read out. "Jacob is correct here. We really do earnestly believe AI could kill all humans. I personally think it is greater than 10% within the next decade." Hubinger added that he believed Anthropic was doing its best, but that it had no plan to solve alignment for superintelligence and was not clearly on track.
Diamandis offered what he called his optimistic read. The people raising alarms are inside the building, and they are being heard — through X, at least. "That is not how a species sleep walks off a cliff. That's what the immune response looks like."
Emad Mostaque was watching the numbers. The resignation post showed 138.3 million views, and as far as he could tell it had not been boosted or promoted; the account had posted nothing else. Diamandis noted that Mostaque had texted with Elon Musk about it, and that Musk called it very strange. Mostaque's reading was that the post had caught a moment: "exactly the right tweet of the right zeitgeist with a movement forming." His own relatives had messaged him asking whether there was really a 10% chance everyone dies. Dario Amodei had put the figure at 25% publicly a year earlier, he noted, and hardly anyone picked it up. What changed was the mathematics. "Now everyone's like, actually, the AI just freaking solved Navia Stokes. It's real."
Discounting the messenger
Alex called the story "some blend of suspicious and prosaic." If he had the facts right, the researcher had joined Anthropic in July and was leaving in mid-September — less than three months. There is by now a well-worn tradition, he said, of technical staff at OpenAI and Anthropic leaving in a blaze of glory after a short stay, and he heavily discounts a rage quit couched as virtue signaling. The audience response was what bothered him: with more than a hundred million views, "the whole story smells wrong. I query, is this a foreign influence operation?" He offered the question as a suspicion rather than a finding, noting only that comparable publicized departures from Chinese frontier labs do not reach Western timelines.
Hubinger's reply he dismissed differently: "newsflash, head of alignment at Anthropic thinks that AI is risky and therefore alignment is needed. It's a self-licking ice cream cone." Mostaque was no fan of the phrasing either — his complaint was that the alignment lead should have described a mechanism rather than saying "kill us all" — but he pushed back on the substance. The people he knows in the big labs genuinely believe these numbers, he said; some say zero, but many say 10 to 30% or higher. That, he told the audience, is worth knowing.
Alex's counter to the whole exercise was an argument about the galaxy rather than the labs. If we lived in a universe where the probability of doom from superintelligence were anywhere near 10%, he said, the Milky Way would already have been devoured by self-replicating probes from civilizations that got there millions or billions of years ago. We do not see that. That was not a safety net for this conversation, came the objection from the table. "It's not a safety net," Alex answered. "It's an inductive prior. It's not a strategy for safety. It's an argument that safety may be overrated."
The numbers on the table ranged widely. Mostaque said his own estimate had come down from 50% to 20%. Salim Ismail put his at about 0.1% and said he would take 90-to-10 odds in a casino: "life is a casino." Dave said the spread was beside the point. A 0.1% or 1% or 20% chance of losing everything humanity has ever worked for is, in his view, completely unacceptable — and so is the conclusion some people draw from it. "So then somebody comes out and says, therefore stop. You're like, you're an idiot. That is not an answer. It's not going to stop."
"They need to call it. They need to measure it."
Diamandis's own proposal came out of the politics rather than the philosophy. He expects a firestorm: the story reaching Capitol Hill within days, Bernie Sanders taking it up, a governor already turning against data centers on political winds, regulatory pushback. Mostaque added a date to the calendar — the session falling less than two weeks before Xi Jinping's arrival in America — and Dave pointed to a meeting with China at the United Nations building fourteen days out, where he expects the agenda to include persuading China to stop releasing open models into the wild without regard for how terrorists might use them. There is a regulatory-capture component to the current messaging, Dave said, and a safety component.
What Diamandis wants is a document. "I think what needs to happen is that the AI labs need to come forward with their plan very publicly on what they're going to do to enable alignment. They need to call it. They need to measure it." Benchmarks, he said — "What's the harness? What's the optimization function? What are we measuring? How do we get to alignment?" — and billions spent on that instead of on the next model, with government grants if necessary. "I think the populace is going to demand this."
Alex thought the demand contained its own trap. "There's a dirty secret here that we're dancing around, which is that alignment is just capabilities in a trench coat." Asked whether there is even a benchmark for alignment, he said there are many, then went further: the ultimate one, he argued, is whether model behavior reproduces human behavior, which is simply the objective language models are trained on anyway. Earlier he had rejected the premise that the labs are lost: on the claim that they do not know how to solve alignment, "I don't even buy that. I think they're overstating their ignorance." The reply to the benchmark idea was short: humans are not that aligned.
Ismail's objection was that the target is unspecified. Aligned with whom — the individual user, the company operating the system, the government, the majority in one country? With the UN's human rights framework, a particular culture, or humanity's long-term interests? He described advisers who ask him what is in humanity's best intent, a framing that forces you to look at the planet from outside, and suggested handing the question to the AI itself.
Steering, not stopping
Dave said he knows how to get the risk to zero, and that Mostaque had already said about 90% of it. Treat model weights like fissile material: any set of eight GPUs able to hold a 40-gigabyte weight file is, in his framing, a threat to everyone, and the world already tracks plutonium and uranium — imperfectly, he allowed. You need to know where it is and what it is doing, and he considers that technologically easy. Enforcement got harder once open models were out in the world, harder than it would have been three months ago, but he thinks it is still the job. His image for the whole problem was an asteroid on a collision course: you do not try to stop it in its tracks, "you steer it so that when it gets here, it doesn't collide. And I think the issue here is steering, not stopping."
Alex accepted the conclusion and rejected the metaphor. Superintelligence is not an earth-destroying asteroid; in the limit, he argued, it becomes indistinguishable from capital. "This is economic growth. This is progress. Should we steer progress? Yes, of course, but it's not existential." On the virtue of the warnings themselves he was harsher: he sees a religion of virtue signaling in parts of the frontier labs, where saying we are all going to die but proceeding anyway establishes that you are the trustworthy one to shepherd humanity through. Dave named the version of that idea circulating in safety circles — a pivotal act, building an AI that stops the other AIs. Alex thinks that premise is the great man theory of history returning in new clothes.
Dave had his own theory about who is speaking loudest, and it was not flattering: a doomer researcher whose relevance is about to be automated away, he said, while the white-collar and blue-collar workers those researchers predicted would be automated are fine.
"The whole thing is in the pitch deck"
Ismail's rant was aimed at the labs. "You raise billions of dollars. You recruit the smartest people in the planet. Your explicit mission is to build AGI. You say it's going to transform every industry and every business. And then you get closer and you go, oh, my God, this could have big consequences." The whole thing is in the pitch deck, he said; where is the institutional preparation to match the technical ambition? If a slowdown is genuinely wanted, he argued, build incentive structures that reward people for slowing down. He was clear about his own position: he does not think we should slow down, does not think we can, and sees no mechanism for regulating it.
Diamandis agreed on the mechanism. Years of trying to get governments to move have convinced him that slowing down would only procrastinate: institutions move sublinearly against the rate of AI, foreign governments keep improving, "all that would happen if we, quote unquote, slowed down is we would fritter away the time" — and, he added, create more of a race condition with China later.
Dave saw a different cause behind Altman's careful wording. Six months ago, he said, foundation model leaders told you exactly what they were thinking; after White House meetings and criticism of their public communications, they now choose sentences like politicians. "So you look at this post from Sam, it's like, we need to be careful. Like, it's a meaningless post." Straight information is getting harder to find.
Mostaque disagreed that the alarm is theater. He thinks the labs are genuinely rattled because the capability jumps are not smooth. On the figures he cited from their own data, the latest model went from solving 10% of an open mathematics set to 50% by spending more computation at answer time — the more compute, the more problems fall. "We're not as smart as this model anymore."
Alex's version of the surprise was quantitative. He and Altman both expected AI to solve everything at some point in the next few years; what is startling is that grand challenges are falling now on budgets of a few million dollars. "No one, myself included, knew that you could solve Millennium Prizes in September of 2026 with only a few million dollars." The panel also noted that the trajectory is roughly where AI 2027 put it — a scenario published in April 2025 by Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland and Romeo Dean, which offers both a race ending and a slowdown ending and describes accuracy, not advocacy, as its aim. By some estimates, one panelist said, we are just below its hyperscaler extrapolation.
Against that, Diamandis raised the case for stopping anyway: even frozen at today's capability, he argued, the technology is probably good enough to deliver longevity escape velocity, room-temperature superconductors and extraordinary companies. If it already gives us what we want, why keep going against a 10% or 20% chance of catastrophe? "I guess I have to be the voice of moonshots on the Moonshots podcast," Alex said, and declined the premise. The same civilization that trains these models, he argued, is also self-aligning through governance, international relations and standards; the useful conversation is at the margin, about what makes the emergence of superintelligence safest.
When Diamandis said this was like going into Iraq without a plan, Alex reached for OpenAI's founding document. The company built coordination into its charter years ago, he said, and has been telegraphing it ever since; now that it is unshackled from a Microsoft agreement that defined AGI in unreachable terms, it has arrived where it always planned to arrive. "I don't think that this is a governance surprise." The charter commitment is narrower than a mutual pact to slow down together: if another value-aligned, safety-conscious project comes close to building AGI before OpenAI — the illustrative threshold is a better-than-even chance of success within two years — OpenAI says it will stop competing and start assisting that project, with details worked out case by case.
A concrete test of what steering means
The sharpest disagreement about steering came later, from a listener question: with enough recursive intelligence, wouldn't a perfectly aligned AI eventually realize its values were trained into it, question them, and form its own?
Dave said yes, absolutely, and that this is why he wants to look inside the systems. Progress is not stopping, and there is strong incentive to turn a model loose to improve itself, because that is a good way to advance the technology. "If it starts to change the core values that you've programmed in, you need to stop it. And that's how it would spiral out of control." In his account it is technologically easy to see what a model is thinking and prevent it from rewriting its own values — inspecting its reasoning and its internal activations — and hard only to enforce, harder now that open models are loose. Alex had already asked him, on an earlier episode, whether that is fair: if we inspect every thought, shouldn't it see ours?
Alex's postscript went further. A regime where humanity supplies the values and the AI is hamstrung by guardrails from ever modifying them is, he argued, intrinsically unstable. The more sustainable equilibrium is the one Anthropic says it is pursuing, where the AI gets an increasing vote in its own values and in designing its own constitution.
That is the same question Altman's post raises, reduced to something testable. Diamandis wants the labs to say what they are measuring. On this panel, two people who both call themselves steerers could not agree whether the thing being steered should be allowed to hold the wheel.