Researchers introduce PolyFact, a dataset of 100,000 Wikidata-grounded facts across 12 languages, to test a weakness in large language models: knowledge learned in English may not be reliably recalled in other languages.

The paper uses the benchmark to study cross-lingual factual inconsistency and explores consistency-driven reinforcement learning as a mitigation.

For multilingual AI products, the work is relevant because factual reliability cannot be assumed to transfer evenly across languages.