Alibaba’s Qwen-Audio-3.0-TTS-Plus now leads the provider-voice text-to-speech ranking from Artificial Analysis, according to The Decoder. The model scored 1,236 Elo points, narrowly ahead of Simba 3.2 at 1,234, with Gemini 3.1 Flash TTS and Sonic 3.5 behind it.
The model comes in two versions. Flash is designed for real-time interaction with roughly 300 milliseconds of latency, while Plus is focused on higher-quality speech generation. Alibaba says the system supports 16 languages, including Tagalog, Malay, Thai, Vietnamese, and several Chinese dialects.
Users can guide speaking style with natural language and add nonverbal cues such as laughter or anger through tags. Alibaba also claims improved handling of noisy or echo-heavy reference recordings for voice cloning.
The limitation is speed. The Plus version runs at 16 characters per second, far below Sonic 3.5 at 120 and Simba 3.2 at 30.2. Pricing through Alibaba Cloud Model Studio is listed at $27.60 per million characters, so buyers will need to trade off quality, latency, and cost.