A new paper proposes replacing fixed safety categories with yes or no questions that operators can define at runtime without retraining the classifier.

Shieldstral, a 3-billion-parameter model from French AI company Mistral, matches models three times its size on standard text safety benchmarks, according to the paper. Mistral says the model also sets a new high score for joint text and image classification.

Runtime rules let operators tailor safety checks

Many guardrail models sort content using fixed taxonomies. The paper's authors, including Mistral co-founder Guillaume Lample, point to two problems with this approach: public safety datasets group risks too differently to support one common taxonomy, and the same rules don't fit every use case. Content suitable for a cybersecurity tool could be harmful on a mental health platform.

Operators write the review criteria in plain language. Shieldstral returns one token, which produces a safety score between zero and one. | Image: Mistral