Mozilla is pushing AI safety work beyond generic benchmarks by arguing that guardrails should be evaluated in the real contexts and languages where people rely on them. At ACM FAccT 2026, the group presented work on testing large-language-model guardrails for refugee and asylum support scenarios across English, Farsi, Arabic, Kurdish-Sorani, and Pashto.
The project used 120 scenario pairs scored by native-speaking evaluators from Respond Crisis Translation on six rights-based criteria. Mozilla says the evaluation exposed recurring problems such as unsafe referrals, missing disclaimers, and stereotyped assumptions, then converted those failures into concrete guardrail policies in English and Farsi.
The practical point is that guardrails are not just invisible filters around a model. They shape what users actually see, especially in high-stakes services. Mozilla argues that open guardrail models and policy-prompt approaches make independent testing possible, but only if the tests reflect specific communities, languages, and failure modes.
The work does not claim a universal safety fix. It shows a process: find local failures, turn them into explicit policies, and test whether those policies actually improve outcomes.