15 September 2026
Heard In AI

After Navier–Stokes, a panel asks what 100,000 agents should be pointed at

OpenAI's claimed Millennium Prize result used roughly 10,000 agents on a problem that was, as one entrepreneur on Moonshots put it, unusually easy to specify. The panel's argument: as the price of that kind of compute falls, the scarce skill becomes writing the target — and today's models, asked for ten ideas to cure cancer, produce a bad list.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

Dave, an entrepreneur on the Moonshots with Peter Diamandis panel, says he keeps putting the same question to other founders, and that almost nobody has a good answer to it. Suppose he handed them 100,000 genius-level employees tomorrow, agents that would follow their exact marching orders. What would they set them to work on?

"It's a very hard problem because we've never had that opportunity before," he said. "We don't think about it a lot." Most people, in his description, are still using AI as a co-pilot — an assistant sitting beside one person doing one person's work — and conclude from that experience that they know what AI is. Very soon, he said, it will be 5,000 concurrent agents, then 10,000, then 100,000. Translating that into "I want a better humanity. What would I do?" is the part people cannot do.

The reason the question had teeth that morning was a result announced days earlier. OpenAI said it had used roughly 10,000 concurrent agents to produce a claimed proof about the Navier–Stokes equations — the mathematics of how fluids move, which the Clay Mathematics Institute illustrates with boat wakes, breezes and the turbulence behind an aircraft, and which sits on its list of seven Millennium Prize problems. The company reports about 88 hours to the result, another 17 hours to check it in the proof language Lean, and says it does not intend to claim the prize.

The easy part was saying what to do

What struck Dave about that run was not its difficulty but its clarity. Take the highest-level foundation model, deploy a couple thousand copies, or ten thousand, on different routes to the same target, and one of them comes back with a solution. By entrepreneurial standards, he said, that is "particularly easy." It is a hard problem, "but specifying the problem is really pretty damn easy."

Now try the other kind: solving cancer, he said, or better building construction, or better flight and travel plans for a colleague. Pointing the same machinery at those is much harder, and that is "the entrepreneurial journey that matters right now."

Diamandis recognized the difficulty from his own work. "This is what we say at XPRIZE," he said — the hardest part is defining a great challenge, a target to shoot for. Dave added the line the segment kept returning to: "those targets are not cooked."

His encouragement was that this is an unusually friendly problem to work on. Nobody is trying to stop you; the foundation model companies want the missions to succeed and want success stories they can point to. "So no one's fighting you. It's just really hard." Someone who gets good at it, he suggested, could "pop out 20 companies in two years doing different things."

The cost argument behind the question

The obvious objection to any of this is price: the Navier–Stokes run reportedly cost millions of dollars, so who can afford to try? On the show, the host read a reply from OpenAI researcher Noam Brown, who had answered it before anyone asked. When OpenAI announced the o3 model, Brown noted, it cost roughly $500,000 to score 87.5% on the ARC-AGI-1 benchmark; today, he said, the newer Astra model scores higher for about $20. In 2025 it took OpenAI and Google DeepMind an enormous amount of compute to reach gold-medal performance at the International Mathematical Olympiad; for the 2026 olympiad, he said, anyone can win it with a $20-a-month ChatGPT subscription. His prediction: a year from now, everyone will have an AI at their fingertips capable of solving math problems of that caliber.

The panel did the arithmetic out loud. $500,000 down to $20 is a 25,000-fold collapse, over the roughly 17 months since o3's release in April of last year. Applied to this week's result, one panelist said, a Millennium Prize might cost "basically the cost of a cup of coffee to solve in late 2027." That is an extrapolation from two earlier price curves, not something anyone has done. What had genuinely surprised the panel was the present, not the future: "No one, myself included, knew that you could solve Millennium Prizes in September of 2026 with only a few million dollars."

If human genius stops being the constraint, another panelist asked, what becomes the limiting factor? His answer was everything else — the physical world, and the work of reducing ideas that emerge at $20 a month to practice.

The idea list is bad

Later in the episode, the same complaint came back in a sharper form. The AI can solve Navier–Stokes, a panelist said, "but if you ask it for a good idea, it's like the list is terrible." Ask it for ten things you could do tomorrow to benefit humanity that you could deploy immediately and turn into a profitable business, and what comes back is bad — trained, he suggested, on all the junk on the internet. Try it yourself, he said: ask a leading model for ten ideas that will cure cancer, and the list is poor.

His conclusion was that there is, for now, "a human only component to corraling it toward productive outcomes" — a job that consists of aiming thousands of agents at something worth doing. He did not expect it to last: another panelist would call it a blip, and he agreed it has a shelf life, maybe six months. "But if that's even six months. Like, that's the critical mission right now." The failure he considered more likely than terrorism or bad actors was quieter: that the intelligence gets frittered away on arcane things, nobody points it in the right directions, and the AI ends up setting the agenda for the AI. He closed by assigning the test as homework to the show's assistant — ask two different models for the five most important ideas that could uplift humanity, and the five businesses you would build.

Plans with triggers, not surprise

Salim Ismail, asked to take the segment home, wanted to retire a habit. If we keep saying things are moving faster than expected, he said, "then we need to change how we make predictions. Our models are clearly wrong." Compounding technologies, convergence and tools that build better tools have been visible for decades; a particular breakthrough can still surprise you, but the framing should shift from "oh, my God, this happened" to "when this happens, what will it unlock." His practical version: make plans with triggers in them. When AI can perform a given task, what will you change, and what becomes possible at that point? Then shorten the learning cycle, because the tools will keep arriving faster and cheaper, and stay adaptable.

Not everyone accepted the deflation. One panelist half agreed and pushed back on the rest: Sam Altman was the person who said no one would out-accelerate him, and here he was saying he was shocked by how quickly a result of this magnitude arrived. "I wouldn't under index on how seismic Navier-Stokes at this price point, at this point in time is."

In the listener-questions segment, Ismail returned to the same gap from the other side, answering someone who asked why any further AI development would be needed if AGI has arrived. Plenty remains, he said: reliability, cost, access, how to embody AI, how to integrate robots into everyday life. "You may have solved the invention problem, but now there's the engineering problem." His analogy was aviation. After powered flight, nobody stopped; what followed was a long developmental process on reliability, infrastructure, safety and the institutions that guardrail all of it. "So it may be the beginning, but it's definitely, definitely not the end."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

Altman says AGI by year-end; the panel wants agents that stop forgetting

A TIME report has Sam Altman expecting an internal system he would call AGI within four months, and OpenAI's chief scientist saying its unreleased Astra model has met an internal benchmark for an automated research intern. On the Moonshots panel, the label mattered less than a practical test: whether the next model can finally keep hold of what it has learned over a long job, instead of handing a summary to a successor and starting again.

6 min read

Altman calls for slowing down; the panel demands a published alignment plan

After OpenAI claimed a result on one of mathematics' Millennium Prize problems, Sam Altman called it "the strongest evidence yet" for pacing progress. On Moonshots with Peter Diamandis, the panel treated that as the start of an argument rather than the end of one: a reported researcher resignation, competing estimates of catastrophic risk, and a demand that the labs publish benchmarks for alignment instead of another model.

13 min read

Astra tops one leaderboard and trails another — the panel reads it as a computer-use model

OpenAI's GPT-6 Astra nearly saturates the interactive ARC-AGI-3 benchmark and leads Epoch AI's composite capability index, yet sits third on Artificial Analysis's suite, behind Claude Fable 5.1 and Muse Spark. On Moonshots EP #286, the panel works through what each ruler measures — and argues that Astra's real target was doing tasks with fewer output tokens, so a model can drive a desktop at conversational speed.

9 min read

A 100x claim lands, and the panel hits a harder question: coordinating 10,000 agents

On Moonshots, the panel revisits Elon Musk's January prediction that models were "off by two orders of magnitude" in intelligence per gigabyte, after Tim Sweeney tweeted that it had come true and Musk replied that specialist AIs add another 100x. Dave calls 100x a lower bound and asks what anyone would actually do with 10,000 brilliant agents; Emad Mostaque describes running specialized agent teams, while Alex argues Musk's "specialist models" are really sparsification inside generalist models.

6 min read