Alibaba's Qwen Audio 3.0 TTS Plus tops Artificial Analysis' Speech Arena
TL;DR
Qwen Audio 3.0 TTS Plus leads the Speech Arena leaderboard with support for 16 languages and natural language style control. It generates speech at 16 characters per second, slower than Sonic 3.5 and Simba 3.2.
What changed
Alibaba released Qwen Audio 3.0 TTS Plus, which now leads Artificial Analysis' Speech Arena leaderboard. The model handles 16 languages and accepts natural language instructions or tags such as [angry] to adjust speaking style.
Specs
- •License proprietary
Why it matters
Vibe Builders gain precise style control through tags or plain language when shaping audio output. Basic Users can access speech in any of the 16 supported languages at a steady 16 characters per second. Developers can benchmark the model directly against Sonic 3.5 and Simba 3.2, which run faster on the same tasks.
What to watch for
Compare output quality against Sonic 3.5 on the Artificial Analysis Speech Arena leaderboard. Developers should run side-by-side latency tests with Simba 3.2 on sample text before integrating the model.
Who this matters for
- Vibe Builders: Use natural language tags like [angry] to control emotional delivery in your audio projects.
Harsh’s take
Qwen Audio 3.0 TTS Plus wins on emotional control and language variety, but the speed is a massive bottleneck. At 16 characters per second, this model is too slow for real-time conversational agents or interactive voice response systems. For builders, the play here is asynchronous content generation where quality and emotional nuance matter more than instant delivery. If you need low-latency voice, stick to Sonic or Simba for now.
by Harsh Desai
About Udio
View the full Udio page →All Udio updatesGo deeper
More from Udio
- FeatureThinking Machines Lab releases Inkling, an open-source model trained on video and audio
Thinking Machines Lab released Inkling, a 975-billion-parameter open source model trained to understand video and audio.