A new arXiv paper proposes a way to measure how much language models lose when the same evidence is presented in different languages. The authors define the Cross-Lingual Comprehension Gap as the reduction in response quality caused by the language of the evidence, while holding the content, question, reference answer, model, and evaluation unit constant.
That distinction matters because many multilingual benchmarks mix together several variables at once. A model may appear capable in one language but fail to preserve the same understanding when evidence shifts to another language.
The work is most relevant for organizations using language models in multilingual settings such as customer support, research, compliance, or education. If a model performs unevenly across languages, translating a workflow from English may create hidden quality gaps. The benchmark is a research contribution, but it highlights a practical point: multilingual support should be tested on equivalent content, not assumed from broad benchmark scores.