Researchers have released Wazobia Eval, a benchmark for testing language models on Nigerian Pidgin beyond basic translation and sentiment analysis. The dataset contains more than 550 manually annotated examples covering emotion understanding, sarcasm detection and culturally grounded reasoning.

Its emotion task uses 16 categories designed to represent registers that conventional positive, negative and neutral labels can miss. The benchmark also supplies standardized tasks and evaluation procedures so that developers can compare models on the same material rather than relying on ad hoc examples.

Nigerian Pidgin is widely spoken but remains underrepresented in mainstream language-model evaluation. That leaves model weaknesses harder to measure and can hide failures that appear only in local humor, implied meaning or culturally specific emotional expression. The authors have made the dataset publicly available through Hugging Face to support reproducible follow-up work.

The release is foundational rather than definitive. A collection of roughly 550 examples cannot represent every region, speaker or usage context, and the paper reports only preliminary pilot evaluations. It provides a common starting point for measurement, not proof that a model performing well will understand Nigerian Pidgin in every real-world setting.