Peter Diamandis came back from a meeting with the leadership of SK hynix and Solidigm — a memory chipmaker and its US storage business — with a conclusion he did not expect: the scarce thing in AI is no longer the graphics processor. "I was so blown away by that meeting at how it's not GPUs; it's actually memory is the rate limiter," he told the Moonshots panel. He posted the thought on X — memory, not compute, is the rate limiter for the agentic era — and said Elon Musk replied "Few realize this," after which the post ran up to 7,000 likes.
The figures he brought to the segment were his own account of what he had been told and read. Memory prices, he said, have climbed 500% in 12 months. Hyperscale cloud companies are reportedly locking in global DRAM production — the working memory chips that sit alongside processors — through 2027. SK hynix's chief executive warned that 2027 would be the worst year in the memory supply industry's history, with demand outstripping production capacity well into the 2030s. Only 2% of the world's memory chips are made in the United States, he added, and while global production rises about 20% a year, AI's appetite for memory is growing at something closer to 200%. His summary: "AI needs memory to think." Every GPU, he said, needs four to six times its cost in memory to function.
Two other items rounded out his briefing. Solidigm, SK hynix's US-based business in flash memory and enterprise solid-state drives, has staged what he called a dramatic turnaround — first-half revenues of $8.6 billion and net margins rising from 3.9% to 47.7%, according to the Nasdaq listing he cited. And Musk's Terafab intends to manufacture memory in-house alongside logic chips, a decision to go vertical across the whole AI manufacturing stack.
Why the demand changed shape
Alex, a regular on the panel, reached first for the pandemic toilet paper shortage as a mental model. Part of that shortage, he said, was not a collapse in production but a change in shape: people stopped going to restaurants and offices, so paper made for enterprise buyers had to be rerouted to households, and the supply chain hiccuped.
Memory is going through its own change of shape. Twenty years ago, he said, a program like Microsoft Word needed comparatively little of it. A frontier model is different. In a transformer — the architecture behind today's large models — every layer has to be loaded into some form of memory so the machine can perform the matrix multiplications that turn an input into an output. Running a model with a trillion parameters therefore has "a very, very different memory footprint" from the software that set the industry's old expectations.
That is the demand side. The supply side, in Alex's account, is where the prices come from. The memory and storage business has historically been boom-and-bust, a pattern Clay Christensen and others have written about, and the people inside it are, in his words, paranoid about when the next bust arrives. That paranoia makes them unwilling to respond elastically: they do not build enough capacity, supply does not rise to meet demand, and prices swing wildly. "Economics 101," he said.
Diamandis put a figure on the paranoia from his own meeting. SK hynix's leadership told him they need to quadruple manufacturing capacity, and that merely doubling it would cost $1.5 trillion — the kind of investment that, in every previous cycle, was followed by a bust. "They're scared of not surviving the next supercycle," he said. Dave, another panel regular, said TSMC had told the same story about chip fabrication plants, which run $20 to $40 billion each: ramp up on the assumption that NVIDIA and Apple will keep ordering, and you may end up overbuilt.
What HBM actually is
The form of memory everyone wants is HBM, high bandwidth memory, described on the panel as roughly 3D: multiple memory layers sitting physically on top of the compute inside a single package. Alex called it "the foothills" of a post-von Neumann architecture — an escape from the decades-old design in which memory and processor are cleanly separated and data shuttles between them, which he compared to the tape of a Turing machine. He remembered DARPA programs searching for a successor 20 years ago. "Well, we found it."
The merger is not finished, and the manufacturers' own description shows where the seams are. In SK hynix's technical explanation of its HBM partnership with TSMC, DRAM dies are stacked above a logic base die and connected vertically through thousands of possible through-silicon vias; the base die is what links the stack to the GPU. The companies proposed using TSMC's advanced logic process for HBM4 base dies to add functionality and support customized performance and efficiency. Their packaging work then connects processor and stacked memory horizontally through an interposer, an arrangement commonly called 2.5D. Stacking memory and putting it beside compute are two different parts of the design, not one accomplished fact.
Dave was blunter about the result. HBM, he said, is "a Rube Goldberg mess": random access memory that is being used to stream sequential files off. "It's so insanely stupid. So it's the most valuable thing in the world right now." His expectation is that better designs arrive soon, helped by AI's speed at inventing things, and that photonic computing and new physics are imminent. Bottlenecks, he argued, do not stop exponentials; they redirect them, with capital flowing into whatever eliminates the constraint.
Alex offered a statistic circulating about HBM's current value: per unit of mass, roughly half its weight in gold. Strip away the packaging, which is most of the weight, and the bare chips are worth considerably more than gold — on the panel's telling, probably the most valuable thing you could carry around in a shoebox. "This is not investment advice."
Etching the weights, and freezing them
Emad Mostaque put memory at about a third of all infrastructure spend today, heading for 50% next year. He does not think that lasts: "the market finds a way." What makes no sense to him is using complicated, expensive HBM to hold static weights — the numbers that define a trained model and do not change while it runs. As weights standardize, he argued, especially for well-defined jobs such as being a decent doctor, the work will move to etching: writing the model into the chip itself.
The panel pointed to companies doing exactly that: Talis, which a speaker said had just been acquired, and Etched, put at a $21 billion valuation. "They're not moving the weights. They're etching them into silicon or into wire on the chip, and then they're massively more efficient because they're not moving around." The same speaker put the potential gain at 100 to 1,000 times, and Dave disclosed that he was talking his own book when Architect Labs was named among the companies pursuing the approach.
The trade-off is in the name. "Once you've etched the weights, then they're frozen." If somebody retrains a better model, customers want to swap to new chips, and the panel's account was that the supply chain is not ready for iteration that fast. The saving Mostaque wants — not paying half a data center build-out for memory alone — comes with the condition that the frontier keeps moving faster than the silicon can be replaced. As he put it, the frontier "can still push it way further than we can imagine."
Whose memory is worth keeping
Diamandis framed the demand as personal: you want your agents to remember everything about you, every interaction, building a world model that understands you, and more agents mean more memory. Alex refined the point toward something he called more scandalous. He does not think individuals carry that much information worth remembering. There is, he argued, an enormous overlap between what a person knows and what the world knows — "it's actually the world knowledge that's what's worth remembering." A model that knows substantially everything about the world, on his account, already knows most of what there is to know about the individual. That pushes the memory burden into shared world knowledge and weights rather than personalized history — and, in the reply that followed, toward a standardized reasoning engine whose world knowledge could be etched.
Who gets to run anything
Earlier in the same episode the panel had already followed the shortage to its consequence for access. If HBM is sold out and GPUs are sold out, one panelist argued, the natural next step is that nobody can run anything except what Anthropic, OpenAI and one or two others serve. Chinese labs can release every open-source model in the world, but "you won't find any place to run it" — and the next generation, at 10 and 20 trillion parameters, needs serious hardware to run the way frontier labs run it. Nobody he meets on the street talks about a universal right to AI today, he said; a year from now, he expects people to be asking what their universal basic right to artificial intelligence is.
The same constraint surfaced again in the audience questions at the end of the episode. Asked whether frontier models can be made more energy efficient instead of building new power, Mostaque answered that necessity is the mother of innovation: "As we run out of energy, as we run out of RAM, you're going to optimize immensely." Asked who holds the real advantage if China leads on power and the United States leads on chips, Dave did not hesitate. Power is a problem — he put the need at 100 gigawatts by the end of the decade against a terawatt already manufactured in the US — but between here and there, he said, it is all about chips. "Every chip, you know, that's why memory is up 5x. So that's the bigger advantage in the short run."