• Sources: primary, discussion
  • Summary: Fireworks states Ember-1 is trained from Kimi K3 to spend fewer tokens on reasoning, reporting a reduction of about 40 percent at comparable benchmark quality. The post states its method rather than the headline alone: per-benchmark cost is computed from public Kimi K3 API pricing, sample counts are given, and the comparison runs against K3 at low, high and max reasoning effort. The post is dated 2026-09-23 and reached this page through a Hacker News submission today. The figures are vendor-reported and unreproduced, and the post states Ember-1 rolls out alongside the base Kimi K3 model as a Research Preview on Fireworks Serverless with two-week serverless access, naming no distribution outside Fireworks.
  • Why it matters: The stated method and the per-effort comparison make the cost claim checkable, while serverless-only access for two weeks makes it a result to verify rather than a model to depend on.

send feedback on this story