• Sources: primary, discussion
  • Summary: HydraFusion chooses per request between a single model, a cascade with a quality gate, and a cross-family critique pass. GitHub's own offline evaluation reports 4.9 percentage points better verified task quality on TerminalBench 2.1 at 67 percent lower estimated cost than Claude Opus 5. Those figures are GitHub's measurements against a competitor's model with no independent reproduction.
  • Why it matters: Runtime model selection moves the cost and quality tradeoff out of configuration and into the request path, and the published numbers cannot yet be checked outside GitHub.
  • Follow-up: Track any independent reproduction of the TerminalBench 2.1 quality and cost figures.

send feedback on this story