A review of 38 primary studies argues that no single existing metric captures whether a real-time voice agent is useful in production. Speech-model papers tend to measure latency, conversation research focuses on turn-taking, and agent benchmarks score task completion, leaving major gaps between a smooth exchange and a correctly completed action.

The authors find that architecture is a deployment constraint rather than a settled contest between end-to-end models and cascades. A chunked pipeline can produce full-duplex behavior—listening and speaking together—even though one cited enterprise account found no fully self-hostable end-to-end system that met its production constraints. Evaluation is also moving toward checking backend state, such as whether a reservation actually exists, instead of accepting the agent’s spoken claim.

The review proposes TRG reporting: timing, recovery after interruptions or errors, and grounded, state-verified outcomes, with an additional dimension for multi-person settings. That last case matters because an agent must decide not only when to speak but which participant is permitted to hear particular information. TRG is a proposed standard, not yet a broadly validated benchmark. It gives teams a practical checklist for testing the conversation and the completed transaction together rather than optimizing natural-sounding audio in isolation.