16 September 2026
Heard In AI

Tag

Training data

Articles about Training data from podcasts, articles and papers, with links to the original sources.

Aaron Levie's question for AI memory: what belongs in the weights?

On Training Data, Box CEO Aaron Levie was asked where enterprise AI memory is heading — retrieval, or models whose weights absorb a company's knowledge. His answer started with a lawyer who can see five matters and whose access changes daily, and ended with a wish for a rubric deciding what gets baked in and what stays a lookup.

6 min read

Better data beat better architecture — but the panel split on its shelf life

A Moonshots panel unpacks Dwarkesh Patel and Jerry Han's experiment, which found that improvements in training data delivered a 12-fold compute-efficiency gain between 2019 and 2025 against 3.7-fold for architectures and training recipes — at small scale, on easy benchmarks. The panel then splits over whether a company's proprietary data is a durable advantage, with a $32 billion data-subsidiary valuation on one side and the fate of BloombergGPT on the other.

7 min read

What changes when nine billion DNA variants are precomputed

Google DeepMind's AlphaGenome Atlas stores predicted molecular effects for every possible single-letter change in the human genome. On Moonshots, the panel called it the "bulk solution" to variant effect prediction, proposed it as a naming system that could help people with the same rare mutation find each other — and argued about why it is not really a lookup table.

5 min read

What OpenAI's 10,000 agents actually proved about fluid flow

OpenAI said on 8 September that an internal model, running roughly 10,000 agents for 88 hours, produced a forced blowup construction for the Navier–Stokes equations and a machine-checked proof of it. On Moonshots with Peter Diamandis, the panel worked through what the result is — a statement about idealized fluids, not a device — what it cost, and why the credit for it was contested within hours.

7 min read

World Labs' Atlas rebuilds a place from photos, and imagines the rest

World Labs released Atlas, a model that generates video along a camera path the user designs and rebuilds scenes from a handful of photographs. Its own garden example shows the seam: one photo leaves the surrounding buildings invented, while more photos pin them down. On Moonshots, the panel worked through what Gaussian splats are and why the approach might matter for robots, games and planning a vacation.

6 min read

OpenAI's Cursor cutoff and two theories about what it is really for

OpenAI has proposed ending the agreement that supplies its models to Cursor, now owned by SpaceX, on 12 November. On the Moonshots panel, one guest read the move as OpenAI betting on its own enterprise stack; another argued the real prize is reasoning traces — the working a model shows while solving a problem. Both explanations lead to the same awkward conclusion: Elon Musk and Anthropic now need each other.

7 min read

Architect Labs' AI-designed chip is running on an FPGA; the 3.4× claim is a projection

On Moonshots, the panel played a launch video for Redwood, an accelerator that Palo Alto startup Architect Labs says its AI designed end to end from a specification written by two architects. The company's paper reports two weeks to verified design and FPGA deployment, with a small language model running in a third week — while the headline 3.4-times efficiency figure comes from a projected Samsung 8-nanometer chip that has not been built. The panel, who disclosed they are investors and an advisor, argued the real story is a "designless" company and recursive self-improvement at the chip layer.

6 min read

A virtual cell that remembers what you did to it

GenBio AI's AIDO Cell simulates a human cell that holds its state across a sequence of interventions, and the Moonshots panel watched a demo and began sketching the end of medicine. The article explains what the simulator does today — prioritizing experiments in two prototype cell lines, with laboratory validation of novel predictions still underway — and separates that from the panel's proposals: an AlphaGo-style search from diseased to healthy cells, open public biology data, and frontier labs paying for all of it.

7 min read

Why Graylin says distillation cannot explain all of China’s AI gains

Asked about allegations that Chinese labs extracted capabilities from Claude, Alvin Graylin argued that access to another model’s answers cannot explain every engineering advance. The Moonshots exchange turned on three distinctions: legitimate distillation versus prohibited extraction, query bills versus development costs, and learning from outputs versus improving the machinery behind them.

6 min read