16 September 2026
Heard In AI

The prediction task may outlive the transformer, a Moonshots panel argues

Asked what comes after large language models, Alex told a caller on Moonshots with Peter Diamandis to separate two things people usually merge: the job of predicting the next piece of text, which he thinks has "effectively infinite longevity," and the transformer machinery doing it, which he says is already being swapped out part by part. Dave added his own forecast that the chips underneath will move to photonics within 18 months to two years.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

A caller on the audience-question episode of Moonshots with Peter Diamandis opened with a shout-out to his wife, Melanie, who had pushed him over and over to start following the panel, and then asked something longer-term. Every technology so far, he said, follows an S-curve: slow at first, then steep, then flattening. What did the panel see on the horizon as a fundamentally new approach after large language models?

Alex took the question and started by splitting it. "First, I would distinguish between LLMs and Transformers," he said. The two words often get used as if they named one thing, and in his view the answer depends entirely on which one you mean.

The job, the machine, and the chip

Three separate layers sit inside the phrase "large language model," and the panel's answer works through all three.

The first is the training objective: the job the system is given while it learns. For an LLM, Alex said, that job is the language modeling task, "which is essentially predicting next words." The model sees text with something missing and has to guess what belongs there. Nobody has to label the data, because the answer is already in the text — this is what makes the task self-supervised.

The second layer is the architecture: the arrangement of mathematical components that carries out the job. The transformer is one such arrangement, introduced in the 2017 paper Attention Is All You Need.

The third is the hardware the whole thing runs on — today, racks of electronic chips.

Alex's argument is that the first layer may never need replacing, while the second is already being replaced. The self-supervised task of "next word, next token, next bit prediction," he said, "has effectively infinite longevity to it. You can scale that task, the self-supervised task, from here to solving everything." He was not sure anything more was needed for AI than predicting the next bit, or, more generally, the missing bit. You could generalize beyond it, he said, but he doubted you had to.

Transformers he treated very differently. They are, as he has said on the show before, being "ship of Theseus style replaced bit by bit" — planks swapped one at a time until little of the original vessel is left. Look at the frontier models today, whether Chinese open-weight or American closed-weight releases, and in terms of their raw pieces they are mostly different from what Attention Is All You Need first came up with. On that layer, he said, "we're already up some sort of sigmoid curve" — the S-curve the caller had asked about.

Which planks are staying

Dave, prompted to open the door further, went through the transformer part by part, and mostly expected its pieces to persist.

The MLP — the multilayer perceptron, also called the feed-forward part, a plain stack of numerical layers that transforms each position in the sequence — is, he said, "identical to what existed 20 or 30 years ago." He called it a pretty good bet that it stays.

Next, the residual connection, which he described as "the residual superhighway": rather than each layer building a fresh representation from scratch, the layer nudges a running vector of numbers that flows straight through the network. "You're just modifying a vector instead of recreating it every layer." That, too, he expects to survive.

The genuinely new ingredient is attention, the mechanism that lets each position in the sequence draw on information from other positions. Here Dave offered an image. "The more you dig in to attention, it's more like a notebook keeping notes than it is like a brain function." As a concept it may get radically restructured, he said, but something will remain in that role: a "knowledge accumulation notebook" sitting behind or beside the MLP at each layer.

The 2017 paper set out a narrower result than the forecasts now built on top of it. Vaswani and colleagues combined multi-head attention, position-wise feed-forward networks, residual connections and positional information, and tested the result on machine translation and English constituency parsing. Training used roughly 4.5 million English–German sentence pairs and 36 million English–French pairs; the larger model trained for three and a half days on eight P100 GPUs and reached 28.4 BLEU on the English–German newstest2014 set, a measure of overlap with reference translations rather than a percentage of correct answers. The authors found learned and sinusoidal ways of encoding position performed almost identically, and that piling on too many attention heads hurt quality. The parsing work also went well, though the model did not beat every comparator. Their conclusion was that attention-based translation trained faster, and they proposed extending the design beyond text.

The layer Dave expects to move fastest

Having kept the architecture mostly intact, Dave put the change underneath it. He said he is "100% sure" the compute substrate will move to photonics — computing carried out with light rather than electrons. After that, he said, physics itself gets reinvented by the current AI, and the substrate may move quicker still to something beyond photonics. The structures above would be the same; they would simply be "running millions of times faster." At that point, echoing Alex, he expected the systems to think of everything else after that and explain it to us.

His timeline for all of this: "in the next really 18 months to two years is my guess." He closed by saying the panel could have talked for hours. "It's just wicked exciting."

Why a prediction machine is not necessarily a parrot

Earlier in the same session, another questioner had raised the objection that hangs over Alex's first claim: if these things are prediction machines, where does originality come from? The pitfall he falls into, he said, is that the LLMs are prediction machines, when what interests him is the emergent "ahas" off the beaten path that spark creativity. A panelist replied that an episode recorded that day, due out the next, argued that AIs are probably being inventive.

Alex called the framing a misconception. Reasoning models, he said, should not be characterized as prediction machines or stochastic parrots — the label for a system that recombines its training text without understanding it. They have been so thoroughly reinforcement fine-tuned on reasoning challenges that calling them prediction machines is, in his words, "tantamount to saying basically that they're incapable of novelty or insight, which is most definitively not true at this point."

That is the hinge of his answer to the S-curve question. Keeping next-bit prediction as the training objective does not, in his account, cap what a system can do, because what gets built around that objective — the architecture, the reinforcement training, and, if Dave is right, the chips — keeps changing underneath it.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

Astra tops one leaderboard and trails another — the panel reads it as a computer-use model

OpenAI's GPT-6 Astra nearly saturates the interactive ARC-AGI-3 benchmark and leads Epoch AI's composite capability index, yet sits third on Artificial Analysis's suite, behind Claude Fable 5.1 and Muse Spark. On Moonshots EP #286, the panel works through what each ruler measures — and argues that Astra's real target was doing tasks with fewer output tokens, so a model can drive a desktop at conversational speed.

9 min read

Grok 4.6 closes the gap—and the panel asks what would take it ahead

xAI’s August 12 release puts Grok 4.6 alongside GPT-5.6 Sol Max in its launch benchmark table, with pricing aimed at sustained agent work. The Moonshots panel’s debate was about the next step: whether training on other models’ reasoning can only help a challenger catch up, and what computing infrastructure it takes to move beyond that.

6 min read

Jiang argues broken trade would turn AI into a control grid, not AGI

On The Diary of a CEO, Professor Jiang held up a semiconductor to argue that technology is "specialization times globalization" — chips designed in California, printed by Dutch machines, made in Taiwan, assembled in China. If that trade fractures, he says, AI will not leap to general intelligence; it will be repurposed to watch people instead. The host offered a different reason to expect the same surveillance.

5 min read

Better data beat better architecture — but the panel split on its shelf life

A Moonshots panel unpacks Dwarkesh Patel and Jerry Han's experiment, which found that improvements in training data delivered a 12-fold compute-efficiency gain between 2019 and 2025 against 3.7-fold for architectures and training recipes — at small scale, on easy benchmarks. The panel then splits over whether a company's proprietary data is a durable advantage, with a $32 billion data-subsidiary valuation on one side and the fate of BloombergGPT on the other.

7 min read