Apple researchers report that explicitly teaching a bilingual speech model to distinguish languages can reduce the performance penalty that often comes with multilingual training. The controlled study used English and French versions of HuBERT, a self-supervised system that learns speech representations from unlabeled audio.
The team tested two interventions: an auxiliary classifier that identifies the language and separate clustering targets for each language. Introducing language discrimination during the first training iteration produced the strongest gains. Later or repeated interventions helped less and caused the model’s representations to separate more sharply by language.
On continuous phonetic discrimination, the bilingual baseline’s error rate fell from 11.6% to 10.4%, compared with 10.8% for the monolingual reference. A lexical score rose from 52.1% to 56.7%, although the monolingual model remained higher at 58.5%. On one prosody measure, performance increased from 68.9% to 72.9%, slightly above the 72.6% monolingual result.
The findings suggest that a multilingual model benefits from sharing patterns across languages but also needs an early signal about which language it is hearing. The experiment covers only a controlled English-French setting, so it does not establish that the same technique will scale unchanged to many languages, data imbalances or production speech-recognition systems.