Mistral introduces Shieldstral to provide lightweight policy-aware moderation for AI models
French artificial intelligence startup Mistral AI SAS today introduced a lightweight multimodal safety artificial intelligence open-weight model that can classify outputs for AI models that outperforms other large language models up to seven times its size, setting a new standard for moderation.
The new model, named Shieldstral, allows developers to write policies in natural language questions at runtime, and the model returns a safety score.
It requires no specialized retraining and stores positive capabilities for both text and images. It also provides a verdict in the form of a single token: a “yes” or a “no,” making the result completely unambiguous.
According to the company, the model has extremely strong text safety, matching or outperforming models that outweigh it by over seven times across diverse safety benchmarks, ranking an overall average safety score of 84.9%. It also achieves an average multimodal image safety benchmark of 83.8%, outperforming all evaluated baselines.








