13 September 2026
Heard In AI

Astra’s robot-arm gains stop short of precision

RoboCurve reports that GPT-6 Astra placed a block in a bowl in 19 of 20 trials, up from 8 of 20 for its predecessor. But it completed only two puzzle insertions—the same as the older model—in a 120-trial evaluation that separates basic manipulation gains from reliable precision.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

A listener to Moonshots with Peter Diamandis asked what would happen if GPT-6 were put inside a humanoid robot—and which benchmarks should judge it. A panelist addressed as Alex pointed to something smaller: a robotic arm. In the podcast discussion, he described RoboCurve’s tests as showing a “near 100% completion rate” on tasks such as moving blocks.

RoboCurve’s evaluation, published on 4 September 2026, gives that claim a precise boundary. GPT-6 Astra placed a block in a bowl in 19 of 20 trials. On the other task, fitting a puzzle piece into its slot, it succeeded just twice in 20 attempts.

What the test involved

RoboCurve compared GPT-6 Astra, Fable 5.1 and Fable 5 using the same YAM robot arms and the same Inspect Robots policy. The arms supplied the physical machinery; the policy supplied the control procedure used with each model. In robotics, a policy determines how a system acts on its observations.

Alex described the model as receiving camera frames and having access to controls for moving the arm. A generalist model—one designed for a broad range of tasks rather than solely for robotics—could therefore use visual information to direct physical actions, instead of merely describing what it saw.

Each model attempted both tasks twenty times, producing 120 recorded trials. Human graders scored attempts on a staged scale running from zero, for no purposeful approach, to four, for successful placement. Intermediate scores captured progress toward the goal without counting it as completion. RoboCurve published the individual trials, keeping partial progress separate from getting the job done.

The task it nearly cleared

On block placement, the progression was steep: Fable 5 succeeded once in 20 attempts, Fable 5.1 eight times, and Astra nineteen times. For Astra’s block task, the report gives a mean run time of 2.5 minutes and an estimated cost of $0.94 per run.

That is the improvement behind Alex’s enthusiasm. Discussing the relationship between cost and completion, he described a trajectory in which costs were falling while capabilities rose. He suggested that robotic manipulation, “at least for some version of robotic manipulation by generalist model,” was approaching saturation: a point where a test offers little room for further gains.

For block placement in this setup, Astra was indeed close to the test’s ceiling. But the second task left much more room.

The precision gap

On puzzle insertion, Astra and Fable 5.1 each completed only two of twenty trials. Fable 5 completed none. The large improvement in block placement did not carry over to full completion of the precision task.

Placing an object in an open bowl and fitting a piece into a matching slot make different demands. Insertion requires the piece to line up with the opening closely enough to fit. The evaluation shows a gap between those outcomes; it does not establish which part of the model-and-control system caused the insertion failures.

Why the panel sees a broader contest

Alex speculated that generalist models would soon move beyond putting blocks into containers and take on more general physical tasks. He contrasted that prospect with vision-language-action models, or VLAs: systems designed to connect visual observations and language instructions to robot actions.

His expectation is that increasingly capable generalist models could challenge those specialized approaches. RoboCurve’s comparison, however, tested three generalist models on two tasks—not generalist models against specialist robotics systems. Nor did its robot arms test a humanoid’s ability to operate reliably across varied activities.

The next hurdle is visible without leaving the benchmark. Astra turned eight successful block placements into nineteen. On the puzzle task, eighteen of its twenty attempts still ended without a successful insertion.

Share this article

Go to the original

Sources & further reading

  1. 01

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

Better data beat model recipes in a small-scale training test

A controlled comparison of 2019–2025 datasets and training recipes reported compute-efficiency gains of 12-fold from data improvements versus 3.7-fold from model recipes. The Moonshots panel explored the business opportunity—and used BloombergGPT’s reportedly short-lived advantage to question whether owning unique data is enough.

5 min read