- Sources: arXiv abstract, HN discussion
- Summary: The paper was submitted to arXiv on 2026-07-31, three weeks before this digest. It states that a diffusion decoder can be obtained by fine-tuning an existing autoregressive model rather than training one, using under 10 percent of the starting model's training token budget, and that the result keeps thinking mode, multimodal inputs and long contexts. It refines blocks of 256 tokens in parallel and reports around 20 tokens per forward pass and roughly 1,500 output tokens per second on a single H100. The paper also states the model still generates autoregressively with minor degradation. Every figure is from the abstract and no independent reproduction was located.
- Why it matters: Retaining autoregressive generation after the fine-tune is what would make a hybrid decoder possible on one set of weights.
send feedback on this story