A listener to Moonshots with Peter Diamandis asked what would happen if GPT-6 were put inside a humanoid robot—and which benchmarks should judge it. A panelist addressed as Alex pointed to something smaller: a robotic arm. In the podcast discussion, he described RoboCurve’s tests as showing a “near 100% completion rate” on tasks such as moving blocks.
RoboCurve’s evaluation, published on 4 September 2026, gives that claim a precise boundary. GPT-6 Astra placed a block in a bowl in 19 of 20 trials. On the other task, fitting a puzzle piece into its slot, it succeeded just twice in 20 attempts.
What the test involved
RoboCurve compared GPT-6 Astra, Fable 5.1 and Fable 5 using the same YAM robot arms and the same Inspect Robots policy. The arms supplied the physical machinery; the policy supplied the control procedure used with each model. In robotics, a policy determines how a system acts on its observations.
Alex described the model as receiving camera frames and having access to controls for moving the arm. A generalist model—one designed for a broad range of tasks rather than solely for robotics—could therefore use visual information to direct physical actions, instead of merely describing what it saw.
Each model attempted both tasks twenty times, producing 120 recorded trials. Human graders scored attempts on a staged scale running from zero, for no purposeful approach, to four, for successful placement. Intermediate scores captured progress toward the goal without counting it as completion. RoboCurve published the individual trials, keeping partial progress separate from getting the job done.
The task it nearly cleared
On block placement, the progression was steep: Fable 5 succeeded once in 20 attempts, Fable 5.1 eight times, and Astra nineteen times. For Astra’s block task, the report gives a mean run time of 2.5 minutes and an estimated cost of $0.94 per run.
That is the improvement behind Alex’s enthusiasm. Discussing the relationship between cost and completion, he described a trajectory in which costs were falling while capabilities rose. He suggested that robotic manipulation, “at least for some version of robotic manipulation by generalist model,” was approaching saturation: a point where a test offers little room for further gains.
For block placement in this setup, Astra was indeed close to the test’s ceiling. But the second task left much more room.
The precision gap
On puzzle insertion, Astra and Fable 5.1 each completed only two of twenty trials. Fable 5 completed none. The large improvement in block placement did not carry over to full completion of the precision task.
Placing an object in an open bowl and fitting a piece into a matching slot make different demands. Insertion requires the piece to line up with the opening closely enough to fit. The evaluation shows a gap between those outcomes; it does not establish which part of the model-and-control system caused the insertion failures.
Why the panel sees a broader contest
Alex speculated that generalist models would soon move beyond putting blocks into containers and take on more general physical tasks. He contrasted that prospect with vision-language-action models, or VLAs: systems designed to connect visual observations and language instructions to robot actions.
His expectation is that increasingly capable generalist models could challenge those specialized approaches. RoboCurve’s comparison, however, tested three generalist models on two tasks—not generalist models against specialist robotics systems. Nor did its robot arms test a humanoid’s ability to operate reliably across varied activities.
The next hurdle is visible without leaving the benchmark. Astra turned eight successful block placements into nineteen. On the puzzle task, eighteen of its twenty attempts still ended without a successful insertion.