- Sources: model card, card source, discussion
- Summary: The Hugging Face page that carried a countdown placeholder earlier on this date is now a model card with weights and a technical report link, and it bills the model as an experimental preview of the Qwen4 architecture. The card states 125B total parameters with 6B activated, plus 51B of n-gram embedding and 4B of MTP, 48 layers, 512 experts with 10 routed and 1 shared activated, a native context of 262,144 extensible to 1,000,000, and a hidden layout pairing Gated DeltaNet with Qwen Sparse Attention. The license is qwen-community-1.0 rather than a standard permissive license. Every benchmark figure on the card is vendor-supplied and several carry harness caveats in the card's own footnotes.
- Why it matters: The architecture is the substance rather than the scores, because the 51B n-gram embedding is a parameter axis that offloads more cheaply than adding experts, and the micro-block sparse attention is what carries the context claim.
- Follow-up: Track independent evaluations and whether the Qwen4 line adopts the same attention layout.
send feedback on this story