Inception releases Mercury 2.5, a diffusion LLM at 1,107 tokens per second
- Sources: primary, discussion
- Summary: Inception states Mercury 2.5 reaches 1,107 tokens per second with a 260K context window, priced at 0.20 dollars per million input tokens and 0.75 dollars per million output tokens.
- Why it matters: That price and throughput profile targets the repeated supporting calls in an agent pipeline rather than the primary reasoning call.