• Sources: Instella-MoE-16B-A3B-Think on Hugging Face, r/LocalLLaMA
  • Summary: AMD's Instella-MoE-16B-A3B repositories were created on Hugging Face on 2026-07-23 as a set of checkpoints spanning pretraining, midtraining, supervised fine-tuning, DPO, and a reasoning variant. The model card describes a 16B-parameter Mixture-of-Experts with 2.8B active per token, 64 routed experts plus 2 shared with 6 active per token, 27 decoder layers, gated multi-head latent attention, and a 128,896-token vocabulary, trained end to end on AMD Instinct MI300X and MI325X GPUs with AMD's Primus framework. AMD states the release includes the training frameworks, data mixtures, intermediate checkpoints, and inference code. The license is ResearchRAIL, which restricts use to academic and research purposes, so this is not an open-weight release in the commercial sense. The card cites arXiv 2511.10628 from November 2025, so the checkpoints post-date the paper describing them.
  • Why it matters: A GPU vendor publishing a full recipe for a run carried end to end on its own accelerators is the clearest public evidence of how far a non-CUDA training stack gets, and the research-only license bounds who can act on it.
  • Follow-up: Watch for a permissive license, a technical report tied to the checkpoint release, and any independent reproduction of the recipe on MI300X.

send feedback on this story