Alibaba's Qwen-Audio-3.0-TTS Takes the Top Spot in Text-to-Speech Rankings

The Core · TL;DR
- Qwen-Audio-3.0-TTS-Plus tops Artificial Analysis' Text-to-Speech leaderboard with an Elo score of 1,236, beating rival commercial voice models.
- Alibaba's Tongyi Lab ships two hosted variants: Flash for low-latency (~300ms) real-time use and Plus for higher-fidelity generation, both covering 16 languages and 20 Chinese dialects.
- Plus leads on speaker similarity (82.75 avg) while Flash edges it on word/character error rate (3.87 vs 3.96), and both support voice cloning with inline emotion tags like [angry] or [giggles].
- Generation speed is a weak point: Plus runs at just 16 characters per second, far slower than Sonic 3.5 (120 cps) and Simba 3.2 (30.2 cps), and access is hosted-only via Alibaba Cloud Model Studio at $27.60 per million characters.
An Elo score of 1,236 is now the number to beat in synthetic voice generation. That's the mark Alibaba's Qwen-Audio-3.0-TTS-Plus posted to claim first place among provider voices on Artificial Analysis' Text-to-Speech leaderboard, putting Tongyi Lab's newest audio model ahead of a crowded field of commercial rivals.
The release, hosted exclusively through Alibaba Cloud Model Studio rather than distributed as open weights, ships in two variants built for different jobs. Flash targets real-time applications, running at roughly 300 milliseconds of latency, while Plus is tuned for output quality over speed. Both cover 16 languages, including English, Chinese, Japanese, Korean, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Indonesian, Malay, Thai, Vietnamese, and Tagalog, alongside support for 20 regional Chinese dialects, a breadth that positions the model for global deployment rather than a single-market product.
Accuracy Versus Speaker Fidelity
On word and character error rate, a standard proxy for how accurately a system reproduces spoken text, Flash actually edges out its sibling, scoring 3.87 against Plus's 3.96 in multilingual testing. Plus reclaims the lead on speaker similarity, averaging 82.75 across all 16 languages compared to Flash's 80.44, meaning it does a better job preserving the vocal identity of a cloned or reference speaker. That split reflects the two models' design intent: Flash sacrifices a fraction of fidelity for interactive speed, while Plus prioritizes the naturalness and consistency needed for content production.
Voice cloning is a core feature of both tiers, and Tongyi Lab says the system now handles noisy or echo-heavy reference audio more reliably than earlier Qwen audio releases, a practical improvement for anyone working with imperfect source recordings rather than studio-clean samples. Developers can also shape delivery through natural language prompts or inline style tags, such as marking a line with [angry] or [giggles], giving finer control over emotional tone without retraining or fine-tuning.
The Latency Trade-off
Ranking first on quality hasn't come without a cost in raw throughput. Qwen-Audio-3.0-TTS-Plus processes text at about 16 characters per second, which lags well behind competitors like Sonic 3.5 at 120 characters per second and Simba 3.2 at 30.2. For workloads that demand rapid bulk generation, that gap could matter more than leaderboard placement.
Under the hood, the model relies on a 12.5 Hz low-frame-rate speech tokenizer paired with a five-stage progressive training pipeline that blends language-model and flow-matching components, an architecture choice aimed at balancing linguistic accuracy with natural-sounding prosody. It also supports one-pass synthesis of audio clips up to three minutes long and includes vocoder super-resolution for 48 kHz output, useful for production-grade audio without post-processing upscaling.
Access runs through Alibaba Cloud Model Studio at $27.60 per million characters for the Plus tier, positioning it as a premium, hosted-only option for teams that prioritize voice quality and multilingual reach over generation speed or self-hosting flexibility.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
