Mistral's new 3B Shieldstral model checks AI inputs and outputs for safety violations using natural language yes-or-no questions instead of fixed categories. It matches models seven times its size in some benchmarks. Operators can set their own criteria at runtime rather than rely on a third party's category system, and the model can run locally.

Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.

Shieldstral is a 3B-parameter policy-adaptive multimodal safety classifier designed to assess text...