14 September 2026
Heard In AI

A 100x claim lands, and the panel hits a harder question: coordinating 10,000 agents

On Moonshots, the panel revisits Elon Musk's January prediction that models were "off by two orders of magnitude" in intelligence per gigabyte, after Tim Sweeney tweeted that it had come true and Musk replied that specialist AIs add another 100x. Dave calls 100x a lower bound and asks what anyone would actually do with 10,000 brilliant agents; Emad Mostaque describes running specialized agent teams, while Alex argues Musk's "specialist models" are really sparsification inside generalist models.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

The host put a tweet on the screen. Tim Sweeney had written that Elon Musk's January 6th prediction on the Moonshots podcast — 100x gains in intelligence at a fixed model size — "was at the edge of plausibility when he made it. Now it's simply a fact." Musk replied underneath: specialist AIs, trained on a single language or a single area of knowledge, are another 100x on top of that.

Then came the clip itself, from the panel's interview with Musk at the Gigafactory: "I think we're off by two orders of magnitude in terms of the intelligence density per gigabyte." Two orders of magnitude, and, as the exchange in the clip has it, "just arithmetic improvement."

Intelligence density per gigabyte is a rough way of asking how much capability fits into a model of a given size. A 100x gain there means either a far more capable model at the same size, or the same answer produced by something a hundred times smaller and cheaper to run. Neither the tweet nor the clip pointed to a standard test; this was a claim and a reply, and the panel spent the segment arguing about what it would mean.

"That's definitely a lower bound now"

Dave had made the same bet earlier, in the episode where the panel wore Christmas hats and gave predictions for the year ahead. The last eight or ten years had been 10x years, he had said; this one would be a 100x year, minimum. Hearing Musk say it at the Gigafactory a couple of weeks later stuck with him.

Now he thinks the number was conservative: "that's definitely a lower bound," he said, and it is "much more likely a thousand to 10,000x year," once you layer the density gain and the specialization gain on top of each other.

His point was not the arithmetic. It was that nobody has a plan for the result. "If I have five or 10,000 agents, all brilliant, working concurrently toward a goal, how do they work together?" he asked. "It's not an easy problem to figure out. Like, we've wanted this for so long that we kind of take for granted that we'll know how to use it when it arrives. Well, here it is."

He turned it into a thought experiment for listeners: someone hands you 10,000 employees tonight, on short notice, and they arrive tomorrow. Everyone says they would do something amazing with that. "Okay, what?"

Point the swarm at itself

Dave had texted the host the same question about his own 5,000-agent experiment. The reply, the host recounted, was to build a model of everything happening at Link Studios — all the companies, all the employees, all the entrepreneurs — simulate their behavior, the way the panel had seen a billion-agent system do in China, and predict which teams will succeed.

Dave liked it for a reason beyond the specific target. The first thing to do with a swarm, he said, is turn it back on its own framework and ask it the question the panel had just asked: have it work out how it should be working. That is how you stay ahead of the capability, "because it's going up far, far faster than you can manage the individual agents" the way people did last year.

Salim Ismail pushed the same exercise outward. Connect the dots, he said, with what Musk did in training Grok on SpaceX's engineering data: imagine having every engineering breakthrough and experimentation technique SpaceX has developed at your fingertips, then ask what problem you would go after with 100x capability. "It's completely an imagination limitation now," he said. "How big do you dare to go?"

Emad's teams, and a prize measured in tokens

Emad Mostaque took the specialization claim literally and reported from inside it. What Musk is now describing, he said, is a hundredfold gain on a cost-per-parameter basis, and it is already visible in models like DeepSeek Flash — models with perhaps ten billion active parameters or fewer once quantized, which can be tuned for very specific jobs.

He runs one. His Grok bot has teams and sub-teams: several are analyzing things for him right now, and they have access to his codex and his Claude Max subscription among other tools. Highly specialized agents like these, he said, will "be able to do a hundred times the compute at the same price because they're that specialized."

If tokens are what now drive progress, he added, someone should put a prize on them: a quadrillion-token X Prize, awarded as people demonstrate impact worth scaling. Salim suggested a hundred trillion might be the more sensible number.

The conversation then turned to the other direction of flow — not tokens produced, but information taken in. Billion-token context windows are coming imminently, the speaker said, and he was "100% sure of this now based on recent results": enough room for an AI to consider the entire Library of Congress "in one thought chunk," three or four orders of magnitude more than a person holds in a single thought. Which is why, he said, compaction — the way the harnesses on both the OpenAI and Anthropic side cope with finite context windows — "has to go." He called it "the enemy of progress in civilization at this point."

"I'm not buying it"

On Musk's second 100x, Alex dissented. Specialized models, he said, are basically another way of saying sparsification — activating only part of a network for any given task. Frontier labs already do this with mixture-of-experts models, where selective activation routes a question to some parameters and not others. Those are teams of specialists already, built into one model that can still be trained end to end.

So he does not expect a bright future for separate specialist models. If anything the arrow points the other way: rather than a chemistry model and a biology model, he expects sparsified activations of a generalist that can scale down to a small parameter footprint and up to trillions of parameters.

Dave stepped in to reconcile the two for listeners. It sounds like a disagreement with Musk, he said, but it is the same effect: you still get the 100x, because you are using a smaller number of parameters to get the exact same thought out. Musk calls them specialist models, which makes them sound like separate things that never touch. In Alex's version they stay connected — "why would you cut them apart?"

Alex accepted the framing and extended it. He construes Musk's prediction as a prediction about sparsity, at two levels he is tracking. The first is the familiar one: fewer parameters active at any moment inside a single model. The second is teams of agents, since agents working together on a common task are arguably a form of sparsification too. On that, he expects "way more teaming."

Which lands back where Dave started. More teaming is the capability. The 10,000 employees arriving tomorrow are still unassigned, and the only concrete answer on the table was to hand the question to the swarm — model Link Studios, work out which teams win, and let the agents propose how they should be organized.

Share this article

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

What changes when an AI agent gets its own computer

On the Moonshots panel, Peter Diamandis runs a Grok Bot chief of staff called Skippy and Emad Mostaque runs 18 of them across his own machines, installing models and making art. Salim Ismail calls it the move from asking an AI to assigning work; Alex argues the messaging-app interface cannot possibly scale.

6 min read

When one bad idea convinces all 5,000 agents

On Moonshots with Peter Diamandis, an operator described a bad coding idea spreading through a swarm of 5,000 identical AI agents until he intercepts and rewinds them — otherwise, he says, roughly $50,000 of tokens goes into a harebrained plan. The panel set that experience beside a new paper on "mind viruses" that spread between agents through editable memory, and argued about whether "virus" is the right word.

6 min read

Grok 4.6 closes the gap—and the panel asks what would take it ahead

xAI’s August 12 release puts Grok 4.6 alongside GPT-5.6 Sol Max in its launch benchmark table, with pricing aimed at sustained agent work. The Moonshots panel’s debate was about the next step: whether training on other models’ reasoning can only help a challenger catch up, and what computing infrastructure it takes to move beyond that.

6 min read

Memory, not GPUs: the shortage that could redesign AI hardware

On Moonshots, Peter Diamandis reported back from meetings with SK hynix and Solidigm leadership with a claim that memory, not compute, now limits AI. The panel argued that a changed workload and a supplier industry scarred by past busts are pushing prices up faster than factories can respond — and that the fix may be new chip designs, including etching model weights into silicon, rather than simply paying more.

8 min read