Alibaba’s Qwen Audio 3.0 TTS Plus leads speech benchmark by two Elo points
Artificial Analysis ranks Alibaba’s latest TTS model first for provider voices, though its benchmark lead is narrow and its output speed trails rivals.
By Renata Fuchs · Policy Reporter
· 2 min read
Alibaba’s Qwen-Audio-3.0-TTS-Plus is ranked first on Artificial Analysis’s Speech Arena leaderboard for provider text-to-speech voices, with an Elo score of 1,236. The model is priced at $27.60 per million characters through Alibaba Cloud Model Studio, and the narrow result matters for AI product teams weighing voice quality against latency, language coverage and throughput.
Artificial Analysis places Qwen-Audio-3.0-TTS-Plus just ahead of SpeechifyAI’s Simba 3.2, which scored 1,234. Gemini 3.1 Flash TTS follows at 1,214, with Sonic 3.5 at 1,207. A two-point lead in an Elo-style ranking is not a clean separation from the nearest competitor, but it gives Alibaba a visible position in a market where leaderboard placement is increasingly used in vendor evaluation.
Alibaba is offering the Qwen Audio 3.0 text-to-speech system in two variants. Flash is intended for real-time interaction and has about 300 milliseconds of latency, according to Alibaba. Plus is aimed at higher-quality speech output, which is the version reflected in the top Artificial Analysis ranking.
The model supports 16 languages. Alibaba lists coverage that includes Tagalog, Malay, Thai and Vietnamese, as well as several Chinese dialects. That language mix is relevant for developers building customer-service, media, education or assistant products outside the English-first markets where many speech systems are evaluated most heavily.
Alibaba also says users can direct speaking style through natural-language instructions. The system supports nonverbal tags such as audio sample prompts for emotions and sounds, including tags like “[angry]” and “[giggles].” The company also claims the new model is better than earlier versions at using reference recordings for voice cloning when the input audio contains noise or echo. Alibaba did not provide a separate third-party score for that cloning claim in the materials cited by Artificial Analysis.
The trade-off is speed. Artificial Analysis reports Qwen-Audio-3.0-TTS-Plus at 16 characters per second. That is well behind Sonic 3.5 at 120 characters per second and Simba 3.2 at 30.2 characters per second. For batch generation of polished audio, the quality ranking may carry more weight. For interactive systems, call-center agents or consumer assistants, throughput and latency can matter as much as the preferred voice score.
Alibaba’s published pricing of $27.60 per million characters gives buyers at least one clear comparison point. The company did not disclose adoption figures, customer names, revenue impact or infrastructure costs tied to Qwen-Audio-3.0-TTS-Plus. For now, the measurable news is the benchmark position: Alibaba has a top-ranked TTS voice model by Artificial Analysis’s methodology, with a slim lead and a clear speed penalty.
This story draws on original reporting from The Decoder.