• Sources: primary, discussion
  • Summary: Designs are submitted as declarative atopile code rather than drawn in a CAD GUI, then built into a circuit graph and a bill of materials and graded by SPICE runs against measured limits, using real manufacturer parts pushed to worst-case tolerance corners. Cost against a reference bill of materials scores only once the circuit works. Results dated September 1 across the 13 tasks in EEBench V1 give Claude Opus 5 61.6 percent, Grok 4.6 57.1, Claude Fable 5.1 56.4, Claude Fable 5 54.3, Claude Opus 4.8 Max 51.4, GPT-5.5 42.3, and GPT-5.6 Sol 39.4, with no GPT-6 Astra result yet. EEBench is built and funded by the atopile team, which sells larger evaluation suites and simulation-backed training environments to frontier labs while stating it does not sell benchmark scores.
  • Why it matters: Deterministic simulation grading fails designs that build cleanly and still miss the job, such as the published submission whose nominal 22 uF part delivered 11.4 uF of effective capacitance at 4.7 V bias.
  • Follow-up: Watch for EEBench or xAI to reconcile the Grok 4.6 model card's 60.0 percent at xhigh reasoning effort against EEBench's own 57.1.

send feedback on this story