Inworld AI launches Realtime TTS-2, ranking #1 on Controlled Voice Arena at $12.50 per million characters
Inworld AI released Realtime TTS-2, a text-to-speech model that ranks #1 on Artificial Analysis' Controlled Voice Arena with an Elo score of 1,123 and #2 on the Provider Voice Arena with an Elo of 1,252. Priced at $12.50 per million characters, it claims to outperform models costing 20 times more, achieves sub-100ms latency, supports over 100 languages, and accepts plain-text delivery instructions for voice modulation.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page itself is projected from evidence records.
- Current automated evidence projection
Cross-source coverage
Wire timeline
Artificial Analysis unveils top AI voice models on Controlled and Provider leaderboards
Artificial Analysis, an AI benchmarking platform, has released its Controlled Voice and Provider Voice leaderboards, ranking top AI voice models. The announcement, shared via X, directs users to view the rankings and vote for models in the Speech Arena. Additionally, the post highlights sample clips from Inworld Realtime TTS-2, available for exploration in the Speech Explorer. This development provides a comparative overview of AI voice synthesis capabilities, enabling users to evaluate and select models based on performance and provider offerings. The leaderboards and voting platform aim to foster community engagement and transparency in the rapidly evolving AI voice technology space.
Realtime TTS-2 ranks #2 on Provider Voice Arena with Elo score of 1,252
A new AI text-to-speech model, Realtime TTS-2, has achieved the #2 position on the Provider Voice Arena, a competitive ranking platform for voice synthesis models. The model earned an Elo score of 1,252, with a margin of error of plus or minus 18, based on 1,094 arena appearances. This performance places Realtime TTS-2 ahead of notable competitors including Alibaba's Qwen-Audio-3.0-TTS-Plus, which scored 1,241, and Speechify Simba 3.2, which scored 1,240. However, it remains behind the current leader, Cartesia Sonic 3.6, which holds the top spot with an Elo score of 1,282. The ranking highlights the rapid advancements and competitive landscape in the field of AI-powered voice generation technology.
Inworld's Realtime TTS-2 tops Controlled Voice Arena, ranks second in Provider Voice Arena behind Cartesia Sonic 3.6
Inworld AI has released Realtime TTS-2, a new text-to-speech model that supports over 100 languages including English, Hindi, Spanish, French, German, Chinese, and Japanese. The model can detect language automatically or accept explicit language settings, and supports multiple languages in a single request. It also allows plain-text delivery instructions. According to Artificial Analysis rankings, Realtime TTS-2 achieved the top spot on the Controlled Voice Arena with an Elo score of 1,123, just 4 points ahead of Cartesia Sonic 3.6. On the Provider Voice Arena, it ranks second with an Elo of 1,252, behind Sonic 3.6 at 1,282. The model processes 106 characters per second and is priced at $20.83 per 1M characters, making it cheaper than Sonic 3.6 ($49), Eleven v3 ($100), and Qwen-Audio-3.0-TTS-Plus ($27.59), but more expensive than Speechify Simba 3.2 ($10).
Show 2 older updatesHide older updates
Inworld AI launches Realtime TTS-2 voice agent at $12.50 per million characters
Inworld AI has released Realtime TTS-2, a new text-to-speech model for voice agents, as announced on X by user DataChaz. The model is priced at $12.50 per million characters, claims to outperform models costing 20 times more, and achieves sub-100ms latency. It also features built-in multi-turn context handling. The announcement highlights the model's ranking as #1 on Artificial Analysis, a platform that benchmarks AI models. This release represents a significant shift in voice agent architecture, offering high performance at a fraction of the cost of competing solutions.
Inworld's Realtime TTS-2 goes GA, ranks #1 in Artificial Analysis Voice Arena
Inworld has announced the general availability of its Realtime TTS-2, a text-to-speech model. The product currently holds the top position, ranking #1, in Artificial Analysis' Controlled Voice Arena. Developers are able to embed delivery instructions directly into the text, ranging from short emotion tags to full sentences describing the desired tone. This release marks a significant milestone for Inworld in the competitive AI voice synthesis market.