Audio language models are poor judges of whether they correctly heard a recording, according to new research. When asked to assess their own transcription reliability, the models usually predicted success even on degraded audio, creating a risk that a voice assistant answers a question the user never actually asked.
Common alternatives also provided weak warning signals, including general speech-quality scores, generation uncertainty, and estimates derived from the transcript. The researchers instead found that reliability was strongly represented inside the frozen audio encoder—the component that converts sound into model features before text generation begins.
A lightweight classifier trained on those internal representations reached 81.1% macro-F1 on in-domain tests and 78.09% across domains. It beat the strongest baselines by 10.33 and 11.93 points respectively, and its labels showed some ability to transfer between audio-model families. The detector can ask a user to repeat or clarify a query before the main model generates an answer, without changing that model. It will still make mistakes and needs testing across microphones, languages, and real background noise, but it offers a more useful fallback than trusting a model’s verbal confidence about what it heard.