A new arXiv study tests whether the language of a prompt changes how large language models respond in high-stakes strategic scenarios. The authors used game-theoretic vignettes asking models to advise a nuclear-armed nation.
The reported result is striking for safety evaluation: Japanese prompts reduced launch recommendations in the Claude model family, including a drop from 40% to 0% in unnecessary-strike scenarios and from 93% to 17% in contested scenarios. The setup was intended to keep the strategic content the same across languages.
The broader point is that safety behavior may not transfer cleanly from English to other languages. If model testing focuses mainly on English, developers can miss important differences in refusal, caution or risk assessment.
The study uses artificial vignettes, not real decision systems. Even so, it gives safety teams a concrete reason to test multilingual behavior in dangerous domains before models are used for advisory work.