- Sources: announcement, weights, discussion
- Summary: Mistral released Shieldstral 1.0, a 3B multimodal safety classifier under Apache-2.0, with weights on Hugging Face. The moderation policy is supplied at inference time rather than baked into the weights. Mistral states the model runs on a single 16GB GPU. Mistral also claims the model matches or outperforms open guard models up to seven times its size and sets a new state of the art on multimodal moderation. Those are Mistral's own figures, which no third party has reproduced, so they are carried as vendor claims and not as results.
- Why it matters: Re-targeting a guardrail to a new deployment context becomes a prompt change rather than a retraining job, and a classifier that fits one 16GB GPU can sit inline per request instead of behind a moderation API.
send feedback on this story