Alibaba's Qwen Audio 3.0 TTS Plus tops text-to-speech rankings
Alibaba's Qwen Audio 3.0 TTS Plus has topped the Speech Arena leaderboard by Artificial Analysis. The model supports 16 languages and allows users to control speaking style via natural language or tags like [angry], but it is significantly slower than rivals, generating only 16 characters per second.
Why it matters: This model sets a new quality benchmark in text-to-speech, but its slow speed highlights the trade-off between quality and latency in AI voice generation.
Full story at: The Decoder ↗