• Sources: Ornith-1.5 release page, HN discussion, r/LocalLLaMA thread
  • Summary: Ornith released a 397B mixture-of-experts model, a 35B mixture-of-experts model, and a 9B dense model, with training built on a self-improvement loop. Ornith reports 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE for the 397B model, against Claude Opus 4.8 at 85.0 and 59.0, so it claims parity on one and trails on the other. The 9B dense model ships a quantized Mobile variant for iPhone and Android and is reported at 70.6 on SWE-bench Verified. Every figure is the vendor's own, averaged over five runs, and the post states harness, temperature, top-p, context window, and anti-reward-hacking safeguards per benchmark, which is why the claim clears the bar for publication. No independent reproduction was located.
  • Why it matters: A 9B dense model reporting agentic coding scores at that level puts the workload within reach of edge hardware, subject to independent confirmation.
  • Follow-up: Track an independent reproduction of the Terminal-Bench 2.1 and SWE-bench Verified figures.

send feedback on this story