Mistral AI unveiled Shieldstral, a 3-billion-parameter multimodal safety model designed to moderate AI outputs. The model lets developers write moderation policies in natural language at runtime and returns safety verdicts as single tokens, achieving an average safety score of 84.9% on text benchmarks and 83.8% on image safety while running on a standard 16GB GPU. Mistral released the model under an open-weight Apache 2.0 license.