Alibaba's new video generator, Wan 3.0, will turn a document, a spreadsheet, a slide deck or a web page into a 30-second video in a single pass. "I wonder what kind of a video would get out of a spreadsheet," host Peter Diamandis said on Moonshots EP #284, before reading out the price: around $0.20 per second of generated video, which by his arithmetic works out to something like $20,000 to $60,000 for a 90-minute film. "These prices are beginning to collapse," he said. The launch landed a day after Alibaba announced the largest private follow-on offering by a Hong Kong-listed company, $10.2 billion. "Clearly, they're engineering their stock price," Diamandis said.
The second item was a claim rather than a product. Diamandis relayed a report from Runway's co-founder that video generation now accounts for roughly 70% of all AI token consumption in China, driven by short-form content and robotics, and growing faster than Claude grew in the United States. The figure reached the show second-hand; nobody on the panel presented it as a measured market statistic.
Tokens are the small chunks of text, image or video data that models read and produce, so counting them is a rough way of asking where a country's computing power actually goes. The framing Diamandis offered was that America is "LLM-pilled" and China is "world-model-pilled": large language models predict the next token, while world models predict the next state of reality. That second kind of model, he noted, matters for robotics, autonomous driving, physics simulation and video generation alike.
The revenue-per-token theory
Alex, first of the guests to answer, said he found the split striking and had what he called a unified field theory for it: "American AI labs are revenue maxing and Chinese AI labs are not."
His example was OpenAI's retreat from video. In his account, OpenAI abandoned the video-generation business because Anthropic "ran away with their lunch" by squeezing more revenue out of each token with code generation, a use case better shaped for a text model's compute. Chinese labs, by contrast, are giving their model weights away for nothing. "You don't revenue max by giving your model weights away for free," he said. And if you are not trying to maximize revenue per token, you may as well spend those tokens on video instead of on more economically productive work. American labs, in his telling, are simply too busy with the profitable applications to run that much world-model inference.
Diamandis tried to turn this into a race: which approach gets to AGI, broadly capable artificial intelligence, faster? Alex called it a trick question, because in his view AGI had already arrived with language models, at a point when there were still no good video models. Asked instead which path moves fastest from AGI toward ASI, superintelligence beyond human capability, he rejected that premise too: "We already have ASI." He proposed a replacement question—which advances the capability frontier more effectively in the long term, "whatever is left of the long term," world models or text-based models—and then answered it himself.
His answer was that omnimodal models win: single systems handling text, vision, audio and video together. He pointed to a leading text model he rated highly whose visual and visual-reasoning abilities he considered weak, called that a medium-term impairment of six to twelve months, and said he hoped Anthropic was busy repairing it through acquisitions. Whether the result comes from diffusion transformers or some hybrid that fuses vision, video and audio with the autoregressive text generation used today, he said, "they have to combine one way or another."
Nobody trusts a Chinese model in the enterprise, Dave argued
Dave warned the others they would have to shut him up on this topic. He agreed with Alex, but said there was a story inside the story: Chinese labs cannot sell their models—he named Kimi and Qwen—into enterprise customers "because nobody knows if they can trust it." Video, by contrast, is perfectly profitable and globally saleable. It is, he said, a fantastic way to take a model whose trustworthiness is in question, get it to market and generate large revenue—"no one's worried about security in that use case."
On the underlying race he was less impressed by the architectural distinction. Which approach recursively self-improves first—that is, which produces a system able to improve the research, code and training methods used to build the next system—is "just purely about smarter engineers with plenty of cash and lots of training chips working on the problem," he said. America's "highbrow road" of solving business problems and curing diseases is simply more profitable per token, and that revenue pours back into more training chips, then into 10-trillion- and 20-trillion-parameter models that get distilled down into smaller ones. That loop, in his account, is what kicks off self-improvement.
He credited Demis Hassabis as the first person to think the strategy through, while arguing he could not act on it inside Google because the company was too bureaucratic. Dario Amodei, he said, took the idea and committed to it: forget consumer movies and videos, focus on a model that can improve itself through better code and better training ideas, and productize it along the way if that happens to work out. The Chinese, Dave added, believe exactly the same thing.
Diamandis, meanwhile, was mourning a lost toy. "I miss Sora," he said. "It was so much fun with my kids." Someone on the panel told him it would not be gone for long.
The physical-world counterargument
Salim thought Dave had identified the real fork—recursive self-improvement or world models—and described the two countries as optimizing around different parts of the economy. The United States is tuning its systems for software engineers and knowledge work; China is heading toward manufacturing, commerce, media and machines, robots included.
His expectation is that the divide does not last. "Eventually you're going to converge," he said: winning systems will have language, vision and memory strung together. And when he thinks about where AI makes the biggest difference, it is at the point where it touches the physical world—humanoid or other robots. "I think the intelligence that can model and then manipulate the world and act inside the world is going to win," he said, which would give China's direction a slight advantage.
He left the conclusion conditional. If Anthropic or one of the other labs reaches recursive self-improvement, then Dave is right and that "trumps everything"—an inner loop that swallows a whole lot of other things. "So you could go either way here."