Alibaba’s Qwen team has released five Audio 3.1 models spanning speech recognition, voice generation and real-time conversation. The lineup separates standard and higher-capability transcription and text-to-speech models, then adds a live system that can listen and speak at the same time.
The standard ASR model improves multilingual and dialect recognition and can remove filler words and repetitions. ASR-Next adds timestamps, multi-speaker identification and detection of emotions, ambient audio and machine noise. For output, the TTS model transfers a voice across languages and lets developers control delivery through text instructions covering emotion, speed and style.
TTS-Next combines a language model with diffusion generation to produce voice, sound effects and background audio together. The real-time model supports interruption during a response. Qwen says it can detect a low mood and answer more slowly and empathetically, but that behavior is a company claim and needs testing in real conversations.
Alibaba is also reducing its audio API prices: about 70% for text-to-speech, roughly 85% for real-time use and as much as 95% for speech recognition. Actual savings will depend on the selected model, workload and regional cloud pricing.