Dave Blundin keeps putting a question to entrepreneurs: suppose they had 100,000 genius-level employees tomorrow, all ready to follow exact marching orders. What would they do with them?
In the Moonshots discussion, he uses that thought experiment to challenge what he calls the “copilot mindset”: treating AI as an assistant for work someone already knows how to do. His harder question is how to turn an ambition—better buildings, better travel, better health—into an assignment that a large AI workforce could usefully pursue.
The proposal is that, as AI becomes better and cheaper at searching for solutions, more of the difficult work may move to defining the problem. That means deciding what success looks like, which constraints matter and how anyone will know whether the result works. Blundin sees an entrepreneurial opportunity there, not merely a need to write better prompts.
Where the idea starts: a target worth aiming at
Blundin draws his contrast from the panel’s discussion of a reported AI solution to the Navier–Stokes mathematical problem. In his telling, many model instances could explore different routes toward the same formal target. The mathematics was hard; specifying the assignment was, by entrepreneurial standards, “really pretty damn easy.”
An AI agent is a system that can work through a task in steps, often using software tools. Running many agents in parallel lets them attempt different approaches at the same time. But the number of attempts does not decide what they should be attempting.
A formally stated mathematical problem supplies an unusually clear destination. A proposed proof must establish a particular claim under particular assumptions. Checking a difficult proof can still require substantial work, but the target is not simply “make mathematics better.”
By contrast, “I want to solve cancer” leaves much of the assignment unwritten. Which cancer? In which patients? Is the aim prevention, longer survival or fewer treatment side effects? What evidence would count as success, and what risks would be acceptable?
Peter Diamandis connects that difficulty to his experience with XPRIZE, which organizes competitions around ambitious challenges. “The hardest part is defining a great XPRIZE challenge, a target to shoot for,” he says. Before competitors can search for a solution, someone has to define a finish line they can aim at and judges can assess.
Consider Blundin’s example of better construction. Turning that ambition into an assignment would mean choosing among goals such as lower cost, faster completion and better building performance. It would also mean setting constraints: a cheaper design cannot count as an improvement if it fails the required safety standard. Finally, the assignment needs a verification step—evidence that the proposed improvement survives outside the design document. These are different decisions from asking an AI to generate more designs.
Supporting material: cheaper search, with conditions attached
The panel’s urgency comes partly from its expectation that AI reasoning will become dramatically cheaper. The discussion cites a claim attributed to OpenAI researcher Noam Brown: that an 87.5% score on ARC-AGI-1 once cost roughly $500,000, while a later o3 result scored higher for about $20. The panel describes that as a roughly 25,000-fold cost decline.
ARC-AGI tests a system’s ability to work out unfamiliar visual patterns. The ARC Prize evaluation of o3-preview documents the expensive end of that comparison, with important conditions. On 100 semi-private tasks, the system scored 75.7% with six samples per task and 87.5% with 1,024 samples per task. The evaluation also covered a separate set of 400 public tasks.
The page’s subsequently revised retail estimate puts the high-compute, semi-private run at about $456,000 overall, or $4,560 per task—not nearly half a million dollars for one puzzle. Those estimates use later pricing assumptions. An April 16, 2025 notice also says the released o3 differed from the preview evaluated there.
That evaluation supports comparing accuracy with the resources needed to achieve it; it does not by itself establish the panel’s full historical-to-current price comparison. Nor does a lower price for answering benchmark questions establish that physical experiments or institutional adoption will become cheaper at the same rate.
The discussion itself turns toward this distinction: even if AI can generate valuable ideas cheaply, putting them into practice may become the limiting factor. More candidate treatments still leave the work of testing them. More building designs still leave the work of constructing and assessing buildings.
The engineering work after invention
Salim Ismail approaches the same issue through an audience question: if artificial general intelligence—AI with broadly capable performance—arrives, why would further AI development be needed?
Even granting that milestone, he says, reliability, cost and access remain unfinished. Putting AI into robots and integrating it into everyday life brings further work. “You may have solved the invention problem, but now there’s the engineering problem,” he says. “How do you move this into the world in an effective way?”
His analogy is aviation. Achieving powered flight did not complete the task of making aviation useful. It opened a long engineering process involving reliability, infrastructure, safety, operating arrangements and institutions to oversee it.
For AI, the equivalent distinction is between demonstrating a capability and delivering a service people can depend on. A robot completing a task once leaves questions about repeatability, operating conditions, maintenance and what happens when it fails. Those questions shape the product as much as the original capability does.
Ismail’s analogy adds a second stage to Blundin’s argument. First define a useful target; then define the conditions under which a successful demonstration can become ordinary practice.
Objections and open questions
Blundin’s imagined workforce is a way to expand entrepreneurial ambition, not a demonstrated account of what 100,000 agents can reliably accomplish. Adding agents does not remove the need to divide the work, reconcile conflicting results and check the output. A clear assignment is necessary for his proposal, but it is not sufficient for success.
There is also a difficulty that mathematical targets can obscure: people may disagree about the finish line. A building’s owner, occupants and builder need not value cost, comfort and completion time in the same way. Choosing a measurable outcome does not settle whose priorities it should represent.
And verification can remain expensive even when search becomes cheap. A promising answer may take years to evaluate in practice. AI might help formulate better assignments and design better tests, too, so the discussion leaves open how much specification will remain distinctively human work. It offers an argument for investing in that skill, not evidence that its scarcity is permanent.
Ismail’s practical response is to plan around demonstrated capabilities rather than repeated surprise. “Make plans with triggers in it,” he says: when AI can perform a particular task, what will change?
For an entrepreneur pursuing better construction, that could mean deciding in advance what evidence would justify moving an AI-generated design into a limited real-world trial. The next assignment would then include not just producing the design, but meeting the safety, cost and performance conditions needed to test it.