- Sources: model card, discussion
- Summary: The Hugging Face card for GLM-5.3-Flash is licensed MIT and states 320B total parameters with 18B active, describing the model as the first natively multimodal entry in the GLM-5 series. The stated architecture combines hybrid sparse and linear attention with Manifold-Constrained Hyper-Connections, over a claimed 30T-token multimodal pre-training corpus, and the card publishes serving paths for SGLang, vLLM, TokenSpeed, and KTransformers. The comparisons to GLM-5.2 and to Claude Opus 4.8 are vendor-run, and the card's own footnote states that its HLE figure is measured with tools on the full set under a 300K context management strategy and judged by GPT-5.6-luna rather than the task default. Z.ai's own blog returned an empty body to this run, so the block rests on the model card.
- Why it matters: An MIT license on a 320B-parameter multimodal checkpoint with published vLLM, SGLang, and KTransformers paths makes it deployable without a negotiated license, which is the part the benchmark table does not decide.
send feedback on this story