Frontier Flux Original

Mistral· 3 min read

Mistral open-weights Shieldstral: 3B policy-adaptive safety classifier

Mistral released Shieldstral on 4 August 2026 as Apache 2.0 open weights: a 3B multimodal safety classifier that takes plain-language policies at inference and returns a continuous yes/no safety score. Hub card is mistralai/Shieldstral-1.0-3B. Vendor tables are competitive with much larger open guards on several text axes and lead VLGuard and UnsafeBench at the stated threshold.

ShareXEmail
Abstract cyan shield geometry on charcoal, safety classifier concept

Abstract concept art for Mistral Shieldstral open-weight safety classifier

Frontier Flux / AI-generated

Mistral released **Shieldstral**, a 3B-parameter open-weight multimodal safety classifier, on 4 August 2026. The pitch is policy at inference time: write a plain-language yes/no safety question, pass the content (text, image, or both), and get a continuous safety score from yes/no logits in one pass. Weights are Apache 2.0 on Hugging Face as mistralai/Shieldstral-1.0-3B. The official blog and arXiv technical report are the product primaries.

From @MistralAI on X (4 August 2026 UTC):

"Introducing Shieldstral, Mistral's 3B open-weights model for content safety that can be deployed on-device"

The blog frames moderation as binary question-answering. Under a fixed system judge prompt, each request has three parts: Instruct (context and optional unsafe definition), Query (a single yes/no policy question), and Document (prompt, response, pair, or image with optional text). The model reads only the yes and no logits and softmax-normalizes them into a continuous score, then you threshold (Hub tables use 0.5). Mistral says one checkpoint covers prompt classification, response moderation, refusal detection, and related checks without a fixed harm taxonomy baked into the weights.

Key facts

  • Open weights under **Apache 2.0**; Hub id `mistralai/Shieldstral-1.0-3B` (API lastModified 4 August 2026 UTC). Base: `mistralai/Ministral-3-3B-Base-2512` with a native Pixtral vision encoder (Hub card).
  • Size and deploy cue: ~3.85B BF16 parameters on the Hub safetensors summary; blog targets a single 16GB NVIDIA GPU; X pitches on-device deployment (not a phone-class claim).
  • Training claim (arXiv abstract): data construction over approximately **54.1M** samples, plus a fine-grained policy-adaptability set. Blog: eval samples held out from training.
  • Vendor-reported evals (blog + Hub tables, stated thresholds/reasoning settings): competitive with open guards up to nearly **7x** its size on several text-safety axes (rows mixed vs GPT-OSS-Safeguard-20B, Qwen3Guard-8B, Nemotron, and others). Multimodal F1 at threshold 0.5: **VLGuard 97.7** and **UnsafeBench 81.8** lead those rows; LlavaGuard row is **72.0**, behind LlavaGuard-7B at 81.4 on the available subset.
  • Multilingual list on the card includes en, fr, es, de, it, pt, nl, zh, ja, ko, ar, ru. Card recommends staying in the **32k** training band.

How to read it

Blog, Hub card, and arXiv abstract are the evidence chain. F1 cells and the nearly-7x line are vendor-reported under named settings, not a third-party leaderboard. The differentiator is policy-adaptive QA at inference instead of a retrain. Image safety still rides Mistral's visual data recipe. Independent harnesses and production error rates remain open. Blog also names Mistral an inaugural Open Secure AI Alliance member with NVIDIA; membership is not an audit.

Independence

Frontier Flux is not affiliated with Mistral AI, NVIDIA, or the Open Secure AI Alliance. Social posts are alerts. Official blog, model card, and paper are the primaries. Vendor F1 cells are not our rankings.

ShareXEmail

Free

Get the morning brief

Free daily email of model releases, benchmarks, and lab notes from the last 24 hours across OpenAI, Anthropic, DeepMind, Meta, NVIDIA, and more. Capability-first, not company drama.

Opens Buttondown to finish signup. Confirm by email if asked. No ads. Unsubscribe anytime.

Back to latest · All Originals · Morning brief