• Sources: primary, organization, discussion
  • Summary: Robocurve ran 120 trials on bimanual I2RT YAM arms under the Inspect Robots harness, and on placing a block in a bowl GPT-6 Astra scored 19 of 20 against Claude Fable 5.1's 8 of 20, at 2.5 minutes per trial to 6.8, about 2.1k output tokens per run to 12.9k, and an estimated $0.94 per run to $2.12. On the harder task, inserting a puzzle piece by its center knob into a matching groove, Astra completed 2 of 20 against Fable 5.1's 2 of 20 and stalled at the same final step. Robocurve publishes per-trial scores and transcripts, and its limitations section states the Astra trials ran two days later rather than interleaved, the bowl comparison used a different rig, and grading was operator-judged with the model known.
  • Why it matters: The capability gain is not uniform across task difficulty, because the model that nearly saturates pick-and-place stalls at the same final step on fine insertion as Claude Fable 5.1 does.
  • Follow-up: Track whether an interleaved rerun on a single rig reproduces the insertion result.

send feedback on this story