An independent arXiv study evaluates OpenAI’s Privacy Filter, a 1.5-billion-parameter detector for personally identifiable information, across 42 synthetic benchmarks in 22 languages and five domains.
The results show clear strengths and weaknesses. The model performed well on structured personal information such as emails and phone numbers, and it outperformed Presidio and XLM-RoBERTa on some PII-annotated benchmarks.
The same evaluation found sharp drops when personal data appeared inside narrative prose. Performance also collapsed for some non-Latin scripts, including Arabic and Cyrillic in the reported tests, while XLM-RoBERTa led on several multilingual named-entity benchmarks.
The practical takeaway is that PII filtering should not be treated as a solved checkbox. A detector may work well for regular patterns and still miss culturally variable names, addresses, or privacy-sensitive facts embedded in ordinary text.