15 September 2026
Heard In AI

World Labs' Atlas rebuilds a place from photos, and imagines the rest

World Labs released Atlas, a model that generates video along a camera path the user designs and rebuilds scenes from a handful of photographs. Its own garden example shows the seam: one photo leaves the surrounding buildings invented, while more photos pin them down. On Moonshots, the panel worked through what Gaussian splats are and why the approach might matter for robots, games and planning a vacation.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

Point a camera at a garden and take one photograph. World Labs' new model, Atlas, will let you fly a virtual camera around that garden — but the cottage and the main house at the edge of the frame are not in your photo, so the model invents them. Take more photographs, of the cottage and of the house, and those buildings stop being guesses. World Labs' announcement uses exactly this example to show where reconstruction ends and imagination begins, and says Atlas can take in more than a hundred input images when more fidelity is wanted.

The company announced Atlas on 1 September as a spatial intelligence model. A user specifies where the camera sits and where it points; Atlas generates the matching views and fills in whatever was never photographed with content that looks plausible. The announcement demonstrates a minute-long video at 1440p following a camera path the user designed.

On Moonshots with Peter Diamandis, Diamandis introduced the release by quoting World Labs chief executive Fei-Fei Li calling it "the best camera conditioned world model ever, opening doors for VFX and robotics." What followed was the panel trying to work out, from the documentation, what Atlas is actually made of — and what that would mean if it scales.

What a "world model" is doing here

A world model is a learned representation of a scene or environment that a system can use to predict what it would see, or what would happen, from a position it has not occupied. Video generators already produce moving pictures from text. The difference Atlas advertises is control: you say where the camera goes, and the generated frames follow that path rather than drifting wherever the model likes.

Emad Mostaque, handed the architecture question by Diamandis, called the release fantastic and said it showed worlds and physics sitting inside video models. He mentioned that one of the pre-training leads on Atlas, Chris Wendler, had previously trained his largest model on the Stability cluster and had sent a note about now working at much greater scale. "Bring on the holodeck," Mostaque said. "I think that's what this is."

The panel's answer to why nobody has that experience in their living room this week was compute. In their account, a global shortage of computing capacity and a fivefold increase in RAM prices are absorbing everything available; without it, one panelist said, something like this would already be deployed at home. It would arrive as a premium product, and people pay for premium experiences.

Transparent blobs as a building material

Alex offered the explanation that made the rest of the conversation legible. Reading the documentation, he said, the core idea appears to be to take a diffusion transformer — the hybrid of a diffusion model and a transformer that he says underpins state-of-the-art American video generators — and add one new modality to the mix of text, images and video it learns from: three- and four-dimensional Gaussian splats.

A Gaussian splat, he said, is "basically a transparent blob." An ellipsoid. Stack and layer enough of them and you get "hyper-realistic-looking, traversable 3D scenes" — scenes you can walk a camera through rather than flat frames. His reading is that World Labs is treating those splats, for the first time at scale as far as he knows, as a first-class training modality alongside the pixels of images and the tokens of text. If that is what is happening, the elaborate camera moves stop being mysterious: once a scene is an arrangement of splats in space, sliding a camera around it is a trivial operation.

Alex was careful that this is his inference from the documentation rather than a confirmed description. But he drew out the consequence. Today's video models generally chop each frame into patches of roughly 16 by 16 pixels and treat each patch as a token, the way a language model treats a word. He finds it remarkable that this worked at all: "Let's just take a language transformer and take images and cut them up and pretend each chunk is a word and just blast it through and see if it works." It worked incredibly well, he said, and nobody would have bet on it. If Atlas's approach scales, Gaussian splats — currently an independent strand of work aimed mostly at gaming — could become "a critical new form of token almost for modeling the physical world." Four-dimensional splats, which add motion over time, were also demonstrated.

One panelist pushed the idea outward: the same process, given the data, could be used to build models of subatomic interactions, of astrophysics at relativistic speeds, or of what happens inside a cell — domains where human intuition is poor. That was offered as a possibility rather than a result.

Robots, and choosing a vacation

The panel has argued before that robots will eventually be trained in high-fidelity simulated worlds rather than in the real one. Alex's response was that this is already the present, pointing to the many embodied and world-model companies training systems from watched video. What Atlas might change, on his reading, is the primitive: not patches of images, but splats.

The seam in the garden example still applies. Atlas fills unseen regions with content that looks right, which is not the same as a measured digital twin of a specific place. On the panel, the wider field — with more world models due from other labs, including Black Forest Labs — was described as approximating reality more and more closely on the way to true digital twins; how close any given scene gets depends on how many photographs the model was actually given.

Asked where this lands for ordinary consumers, the panel went past video games to planning. Instead of an AI that hands you a day plan in text, one panelist suggested, you would step into it: "Let's walk through what it would be like to do this or to do that, to play tennis or to go to the Caribbean." Diamandis extended it to three candidate vacation destinations explored as a family before deciding which one to book, and to the ordinary pleasure of returning to a resort you already know — not wandering, not queuing, not signing up for something that turns out to be the wrong thing.

That forecast sat alongside a nearer-term project Diamandis mentioned in passing: a pilot building a digital twin of an organization by combining world models with language models, which he said he would report on in the coming weeks, with compute as the likely constraint.

The segment ended where new vocabulary usually does. "I guess we'll find out how many 3D Gaussian splats it takes to model the water cooler at the office," one panelist said. Another admitted it was not on his bingo card: "Gaussian splats, new word for me."

Share this article

Go to the original

Sources & further reading

  1. 01

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

Why a Moonshots panel thinks China's AI tokens go to video and America's to code

Alibaba's Wan 3.0 and a relayed claim that 70% of Chinese AI token use goes to video sent the Moonshots panel into an argument about money: one guest said American labs chase revenue per token while Chinese labs give their weights away, another said video is the only market that will trust a Chinese model. They ended up disagreeing about whether world models or text models reach self-improving AI first.

6 min read

Higgsfield’s AI feature tests how far filmmaking costs can fall

The Moonshots panel described The Cully Hill Boys as a 110-minute AI feature made by 28 people in four weeks for roughly $2 million, half of it spent on computation. Emad Mostaque expects cheaper video generation to cut that bill sharply. But the film’s published production materials are study-only, and cheaper footage does not resolve performers’ rights or the need for creative direction.

5 min read

A virtual cell that remembers what you did to it

GenBio AI's AIDO Cell simulates a human cell that holds its state across a sequence of interventions, and the Moonshots panel watched a demo and began sketching the end of medicine. The article explains what the simulator does today — prioritizing experiments in two prototype cell lines, with laboratory validation of novel predictions still underway — and separates that from the panel's proposals: an AlphaGo-style search from diseased to healthy cells, open public biology data, and frontier labs paying for all of it.

7 min read

Mostaque wants every child to own a slice of their state's AI company

On the Moonshots panel, Emad Mostaque presented what he calls a Champion: a locally owned intelligence utility, sold to residents at a symbolic $1 pre-money valuation before outside investors arrive, with 10% of the equity set aside in perpetuity for every child under 20. He borrows the structure from his account of how TSMC was capitalized, and argues that as the cost of intelligence falls, the money will sit in robots, deployment engineers and citizen agents. He calls it an idea, not an offering — and it leaves governance, dilution and distribution unsettled.

7 min read