A new arXiv paper argues that voice AI needs different turn-taking behavior for different situations. Current full-duplex systems often apply one timing norm, even though people interrupt, pause, and yield differently in cooperative and competitive tasks.

The authors introduce DuplexGen, a framework for generating dialogues with scenario-adaptive turn-taking. It calibrates language-model predictions against a small set of slot-level human preference annotations instead of relying only on large human-human speech corpora or prompts.

Across six cooperative and competitive tasks, the paper reports systematic differences in what people prefer. DuplexGen aligned more closely with those preferences than uncalibrated prompting or training only on generic conversation data.

A full-duplex model trained on DuplexGen-generated data also showed distinctive, human-preferred behaviors. The finding is practical for voice assistants: natural conversation is not just about speech quality, but about when the AI speaks, waits, interrupts, or backs off.