14 September 2026
Heard In AI

Why a Moonshots panel thinks China's AI tokens go to video and America's to code

Alibaba's Wan 3.0 and a relayed claim that 70% of Chinese AI token use goes to video sent the Moonshots panel into an argument about money: one guest said American labs chase revenue per token while Chinese labs give their weights away, another said video is the only market that will trust a Chinese model. They ended up disagreeing about whether world models or text models reach self-improving AI first.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

Alibaba's new video generator, Wan 3.0, will turn a document, a spreadsheet, a slide deck or a web page into a 30-second video in a single pass. "I wonder what kind of a video would get out of a spreadsheet," host Peter Diamandis said on Moonshots EP #284, before reading out the price: around $0.20 per second of generated video, which by his arithmetic works out to something like $20,000 to $60,000 for a 90-minute film. "These prices are beginning to collapse," he said. The launch landed a day after Alibaba announced the largest private follow-on offering by a Hong Kong-listed company, $10.2 billion. "Clearly, they're engineering their stock price," Diamandis said.

The second item was a claim rather than a product. Diamandis relayed a report from Runway's co-founder that video generation now accounts for roughly 70% of all AI token consumption in China, driven by short-form content and robotics, and growing faster than Claude grew in the United States. The figure reached the show second-hand; nobody on the panel presented it as a measured market statistic.

Tokens are the small chunks of text, image or video data that models read and produce, so counting them is a rough way of asking where a country's computing power actually goes. The framing Diamandis offered was that America is "LLM-pilled" and China is "world-model-pilled": large language models predict the next token, while world models predict the next state of reality. That second kind of model, he noted, matters for robotics, autonomous driving, physics simulation and video generation alike.

The revenue-per-token theory

Alex, first of the guests to answer, said he found the split striking and had what he called a unified field theory for it: "American AI labs are revenue maxing and Chinese AI labs are not."

His example was OpenAI's retreat from video. In his account, OpenAI abandoned the video-generation business because Anthropic "ran away with their lunch" by squeezing more revenue out of each token with code generation, a use case better shaped for a text model's compute. Chinese labs, by contrast, are giving their model weights away for nothing. "You don't revenue max by giving your model weights away for free," he said. And if you are not trying to maximize revenue per token, you may as well spend those tokens on video instead of on more economically productive work. American labs, in his telling, are simply too busy with the profitable applications to run that much world-model inference.

Diamandis tried to turn this into a race: which approach gets to AGI, broadly capable artificial intelligence, faster? Alex called it a trick question, because in his view AGI had already arrived with language models, at a point when there were still no good video models. Asked instead which path moves fastest from AGI toward ASI, superintelligence beyond human capability, he rejected that premise too: "We already have ASI." He proposed a replacement question—which advances the capability frontier more effectively in the long term, "whatever is left of the long term," world models or text-based models—and then answered it himself.

His answer was that omnimodal models win: single systems handling text, vision, audio and video together. He pointed to a leading text model he rated highly whose visual and visual-reasoning abilities he considered weak, called that a medium-term impairment of six to twelve months, and said he hoped Anthropic was busy repairing it through acquisitions. Whether the result comes from diffusion transformers or some hybrid that fuses vision, video and audio with the autoregressive text generation used today, he said, "they have to combine one way or another."

Nobody trusts a Chinese model in the enterprise, Dave argued

Dave warned the others they would have to shut him up on this topic. He agreed with Alex, but said there was a story inside the story: Chinese labs cannot sell their models—he named Kimi and Qwen—into enterprise customers "because nobody knows if they can trust it." Video, by contrast, is perfectly profitable and globally saleable. It is, he said, a fantastic way to take a model whose trustworthiness is in question, get it to market and generate large revenue—"no one's worried about security in that use case."

On the underlying race he was less impressed by the architectural distinction. Which approach recursively self-improves first—that is, which produces a system able to improve the research, code and training methods used to build the next system—is "just purely about smarter engineers with plenty of cash and lots of training chips working on the problem," he said. America's "highbrow road" of solving business problems and curing diseases is simply more profitable per token, and that revenue pours back into more training chips, then into 10-trillion- and 20-trillion-parameter models that get distilled down into smaller ones. That loop, in his account, is what kicks off self-improvement.

He credited Demis Hassabis as the first person to think the strategy through, while arguing he could not act on it inside Google because the company was too bureaucratic. Dario Amodei, he said, took the idea and committed to it: forget consumer movies and videos, focus on a model that can improve itself through better code and better training ideas, and productize it along the way if that happens to work out. The Chinese, Dave added, believe exactly the same thing.

Diamandis, meanwhile, was mourning a lost toy. "I miss Sora," he said. "It was so much fun with my kids." Someone on the panel told him it would not be gone for long.

The physical-world counterargument

Salim thought Dave had identified the real fork—recursive self-improvement or world models—and described the two countries as optimizing around different parts of the economy. The United States is tuning its systems for software engineers and knowledge work; China is heading toward manufacturing, commerce, media and machines, robots included.

His expectation is that the divide does not last. "Eventually you're going to converge," he said: winning systems will have language, vision and memory strung together. And when he thinks about where AI makes the biggest difference, it is at the point where it touches the physical world—humanoid or other robots. "I think the intelligence that can model and then manipulate the world and act inside the world is going to win," he said, which would give China's direction a slight advantage.

He left the conclusion conditional. If Anthropic or one of the other labs reaches recursive self-improvement, then Dave is right and that "trumps everything"—an inner loop that swallows a whole lot of other things. "So you could go either way here."

Share this article

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

Would she pay the real price? Zitron's test for AI adoption

On The Diary of a CEO, critic Ed Zitron praises a chatbot for reading a troubleshooting log and for helping fix his son's Minecraft mod, then argues that neither is worth a trillion dollars. The host counters with his fiancée's one-woman business and his chief of staff's inbox. The argument turns on tokens, subscription rate limits and who is paying the real bill.

7 min read

Altman says economic inertia slowed AI's impact. His podcast panel disputes the cause

On Moonshots with Peter Diamandis, the panel watched Sam Altman explain that he expected GPT-4 to put software businesses up for grabs far sooner than it did, and that the economy's inertia has made the transition "smoother and slower." Salim Ismail blamed institutions that move at a different speed from the technology, Alex pointed instead at the abstraction layers of the economy and prescribed vertical integration, and Emad Mostaque objected that the models simply were not good enough until recently.

6 min read