- Sources: primary, discussion
- Summary: Cognition states SWE-2 is post-trained from Kimi K3, a 2.8 trillion parameter model, and that its reinforcement learning adds 5 to 6 points on many benchmarks over that base. The stated method trains all reasoning-effort levels in one run with a reward of success minus a linear cost penalty, tuned per effort level to the local slope of the base model's Pareto curve, and the reported behavior change is fewer steps, with a first real edit after a median of 18 steps against 48 for SWE-1.7. Every benchmark figure is Cognition's own, and the same table puts SWE-2 at 27.3 percent on Terminal-Bench 4 against 55.8 for Fable 5.1 and 57.9 for GPT-6 Astra, so the cost-frontier claim does not hold on every benchmark it reports.
- Why it matters: A commercial US coding product built on Chinese open weights is a concrete outcome of open-weight releases, and SWE-2 is available today in Devin Desktop and CLI.
send feedback on this story