Mistral has introduced Shieldstral, a 3-billion-parameter open model designed to check AI inputs and outputs for safety violations. The Decoder reports that the model uses natural-language yes-or-no questions instead of a fixed moderation category system.

That design gives operators more control over the rules they apply. Rather than relying only on a provider’s predefined taxonomy, a team can express safety criteria at runtime and run checks against its own policy needs.

The model’s small size is the practical hook. Mistral says Shieldstral matches models around seven times larger in some benchmarks, which could make local or private moderation more feasible for teams that cannot send every request to a hosted service.

Benchmarks are not the same as production safety. The model’s usefulness will depend on how well its yes-or-no judgments hold up across languages, edge cases, and adversarial prompts.