- Sources: primary, discussion
- Summary: The concrete anchor is Gulrajani and Hashimoto's May 2023 measurement that a likelihood-based continuous diffusion model was 64x less training-efficient than an autoregressive baseline, which Dieleman argues was disqualifying while the field was still optimising Chinchilla-style compute against perplexity. The technical section covers embedding strategies from one-hot through pre-trained to jointly learned and the collapse failure mode of the last, the pairing of loss function to unembedding strategy, and why noise schedules matter more here because meaningful corruption of high-dimensional embeddings happens across a narrow band of noise levels. He flags the account as subjective and the causal claims as speculative.
- Why it matters: A DeepMind research scientist writing on his own domain sets out the design choices and failure modes anyone building a continuous diffusion language model has to make first.
send feedback on this story