A new Hugging Face benchmark examines how frontier speech-recognition systems handle code-switched conversations, where a speaker moves between languages in the same interaction.

That matters for voice agents used in customer service, where bilingual speech is common and transcription errors can break downstream reasoning or routing.

The benchmark gives teams a more realistic way to evaluate voice AI than single-language accuracy scores alone.