AWS has published a SageMaker AI tutorial for streaming speech from Qwen3-TTS before the model finishes generating an entire response. The setup uses the company’s vLLM-Omni Deep Learning Container and a persistent bidirectional connection, allowing an application to send text while receiving audio chunks continuously.

Starting playback early reduces the silent delay that makes voice agents, accessibility tools, tutoring applications, and customer-service systems feel unresponsive. The sample includes a Gradio interface and complements an earlier AWS workflow that streamed microphone audio into a Voxtral speech-to-text model. Together, the two patterns cover the input and output sides of a real-time spoken interaction.

vLLM-Omni extends the vLLM serving system beyond text to audio, images, and video, while the AWS container packages tracked releases and adds SageMaker routing middleware. This tutorial is a deployment example, not a latency guarantee: response time and cost will still depend on the selected model, instance, network path, and workload. Developers can clone the sample and reproduce the speech pipeline, but they should benchmark it with their own concurrency and audio-quality requirements before using it in a live service.