Alibaba’s Qwen-Audio-3.1 Voice AI Models Launch with Price Cuts Up to 95%
On September 23, 2026, Alibaba released the Qwen-Audio-3.1 series of voice AI models at the Yunqi Conference, featuring upgrades to ASR, TTS, and real-time interaction models, plus two new models (ASR-Next and TTS-Next). The company cut prices across the entire Qwen-Audio lineup by up to 95% and open-sourced the Qwen-Audio-Agent framework.
IllustrationEditorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary awaiting refresh
Summary awaiting refresh
Cross-source coverage
Reporting timeline
Alibaba Releases Qwen-Audio-3.1 Voice AI Models at 2026 Yunqi Conference, Cuts Prices Up to 95%
On September 23, at the 2026 Yunqi Conference, Alibaba announced the release of its Qwen-Audio-3.1 series of voice AI models, including upgraded versions for speech recognition (ASR), speech synthesis (TTS), and real-time voice interaction (Realtime). The API services are now available on the Qianwen AI platform. To reduce user costs, Alibaba cut prices across the Qwen-Audio voice model line by up to 95%. The ASR model features enhanced multilingual and dialect recognition, contextual understanding, and new native transcription polishing that removes filler words and reorganizes semantics. A next-generation audio understanding model, Qwen-Audio-3.1-ASR-Next, can interpret human emotions, environmental sounds, and mechanical noises, enabling sound description, event localization, and audio Q&A. The TTS-Next model uses a unified generation framework combining language and diffusion models to produce human voice, sound effects, and background audio simultaneously, boosting efficiency for audiobooks, films, podcasts, games, and advertising. The Realtime model improves listening, conversation flow, and empathy, and can proactively invoke Agent tools. Alibaba also open-sourced the Qwen-Audio-Agent framework for developers to build real-time voice agents.
Read sourceQwen Releases Qwen-Audio-3.1 Voice Models, Cuts Prices Across Entire Audio Line
On September 23, Qwen (千问) officially released the Qwen-Audio-3.1 series of voice large models, according to a report from financial news outlet 财联社. The upgrade includes comprehensive improvements to three core models: automatic speech recognition (ASR), text-to-speech (TTS), and real-time voice interaction (Realtime). Additionally, Qwen launched two new models: the audio creation model Qwen-Audio-3.1-TTS-Next and the audio understanding model Qwen-Audio-3.1-ASR-Next. The five new models together form a complete audio capability stack covering 'understanding, generation, interaction, and creation.' To reduce user costs, Qwen has lowered prices across its entire Qwen-Audio voice model line. The price reductions include approximately 70% for TTS, about 85% for Realtime, and up to 95% for ASR.
Qwen Releases Qwen-Audio-3.1 Voice AI Models, Cuts Prices Across the Board
On September 23, Qwen officially launched the Qwen-Audio-3.1 series of voice AI models. The upgrade includes comprehensive improvements to three core models: Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Real-time Voice Interaction (Realtime). Additionally, Qwen introduced two new models: the audio creation model Qwen-Audio-3.1-TTS-Next and the audio understanding model Qwen-Audio-3.1-ASR-Next. The five new models form a complete audio capability stack covering 'understanding, generation, interaction, and creation.' To reduce user costs, Qwen has lowered prices across its entire Qwen-Audio voice model line. The TTS model price has been reduced by approximately 70%, the Realtime model by about 85%, and the ASR model by 95%.
Read sourceShow 8 older updatesHide older updates
Qwen Releases Qwen-Audio-3.1 Voice Models, Cuts Prices Across Entire Audio Line
On September 23, Jin10 Data reported that Qwen (千问) officially released the Qwen-Audio-3.1 series of voice large models. The upgrade includes comprehensive improvements to three core models: automatic speech recognition (ASR), text-to-speech (TTS), and real-time voice interaction (Realtime). Additionally, Qwen launched two new models: the audio creation model Qwen-Audio-3.1-TTS-Next and the audio understanding model Qwen-Audio-3.1-ASR-Next. The five new models form a complete audio capability stack covering 'understanding, generation, interaction, and creation.' To further reduce user costs, Qwen cut prices across its entire Qwen-Audio voice model line. Specifically, TTS prices were reduced by approximately 70%, Realtime by about 85%, and ASR by up to 95%.
Qwen releases Qwen-Audio-3.1 voice models, cuts prices across entire lineup
Qwen has officially released the Qwen-Audio-3.1 series of voice AI models, featuring upgrades to core models for automatic speech recognition (ASR), text-to-speech (TTS), and real-time voice interaction (Realtime). The launch also introduces two new models: the audio creation model Qwen-Audio-3.1-TTS-Next and the audio understanding model Qwen-Audio-3.1-ASR-Next. The five new models form a complete audio capability stack covering understanding, generation, interaction, and creation. To lower user costs, Qwen has reduced prices across all Qwen-Audio voice models, with TTS prices dropping by approximately 70%, Realtime by about 85%, and ASR by 95%. The information was sourced from the科创板日报 (Star Market Daily).
Read sourceQwen Releases Qwen-Audio-3.1 Voice AI Models, Cuts Prices Across Entire Lineup
On September 23, Qwen officially released the Qwen-Audio-3.1 series of voice AI models. The upgrade includes comprehensive improvements to three core models: automatic speech recognition (ASR), text-to-speech (TTS), and real-time voice interaction (Realtime). Additionally, two new models were introduced: the audio creation model Qwen-Audio-3.1-TTS-Next and the audio understanding model Qwen-Audio-3.1-ASR-Next. These five new models form a complete audio capability stack covering understanding, generation, interaction, and creation. To further reduce user costs, Qwen has lowered prices across its entire Qwen-Audio voice model lineup. The TTS model price was reduced by approximately 70%, the Realtime model by about 85%, and the ASR model by 95%.
Read sourceQwen Releases Qwen-Audio-3.1 Voice AI Models, Cuts Prices Across Entire Audio Line
On September 23, Qwen officially released the Qwen-Audio-3.1 series of voice large models, according to a report by People's Financial Information and published by East Money. The upgrade includes comprehensive improvements to three core models: automatic speech recognition (ASR), text-to-speech (TTS), and real-time voice interaction (Realtime). Additionally, Qwen launched two new models: the audio creation model Qwen-Audio-3.1-TTS-Next and the audio understanding model Qwen-Audio-3.1-ASR-Next. The five new models form a complete audio capability stack covering understanding, generation, interaction, and creation. To further reduce user costs, Qwen has lowered prices across its entire Qwen-Audio voice model line. The TTS model price was reduced by approximately 70%, the Realtime model by about 85%, and the ASR model by 95%. The article is sourced from Securities Times.
Read sourceAlibaba releases Qwen-Audio-3.1 voice model series at 2026 Yunqi Conference
At the 2026 Yunqi Conference on September 23, Alibaba released its new Qwen-Audio-3.1 series of voice models. The series includes three main components: an enhanced ASR (Automatic Speech Recognition) model with improved multilingual and dialect recognition, contextual understanding, and a new native transcription polishing capability that removes filler words and reorganizes semantics. A next-generation audio understanding model, Qwen-Audio-3.1-ASR-Next, can understand human emotions, environmental sounds, and mechanical sounds, enabling sound description, event localization, and audio Q&A. For speech synthesis, the Qwen-Audio-3.1-TTS-Next model uses a unified generation framework combining language and diffusion models to generate human voice, sound effects, and background sounds simultaneously. The upgraded Qwen-Audio-3.1-Realtime model features enhanced empathy and can proactively call Agent tools. Alibaba also open-sourced the Qwen-Audio-Agent framework for developers to build real-time voice agents.
Read sourceAlibaba releases Qwen-Audio-3.1 voice AI models with up to 95% price cut
On September 23, Alibaba released the Qwen-Audio-3.1 series of voice AI models, featuring comprehensive upgrades across three model types: automatic speech recognition (ASR), text-to-speech (TTS), and real-time voice interaction (Realtime). The API services for all three models are now available on the Qianwen AI platform. To reduce user costs, Alibaba has cut prices across the entire Qwen-Audio voice model line, with reductions of up to 95%. Specifically, the Qwen-Audio-3.1-ASR model enhances multilingual and dialect recognition, improves contextual understanding, and adds native transcription polishing capabilities that automatically remove filler words and repeated expressions while reorganizing semantics. The Qwen-Audio-3.1-Realtime model can proactively call Agent tools to complete tasks. The article is sourced from Beijing Business Daily.
Read sourceQwen Releases Qwen-Audio-3.1, Cuts Voice Model Prices by Up to 95%
According to a report from First Financial, cited by East Money, Qwen (a major AI model provider) has officially released the Qwen-Audio-3.1 series of voice AI models. The upgrade encompasses three core models: automatic speech recognition (ASR), text-to-speech (TTS), and real-time voice interaction (Realtime). Additionally, Qwen introduced two new models: the audio creation model Qwen-Audio-3.1-TTS-Next and the audio understanding model Qwen-Audio-3.1-ASR-Next. To reduce user costs, Qwen has lowered prices across its entire Qwen-Audio voice model lineup. The TTS model price was reduced by approximately 70%, the Realtime model by about 85%, and the ASR model by 95%. The announcement was made via Qwen's official channels.
Read sourceQwen-Audio Voice Model Series Price Cuts Announced, Up to 95% Reduction
On September 23, Qwen officially released the Qwen-Audio-3.1 series of voice large models and announced price reductions across its entire Qwen-Audio voice model lineup. The price cuts include approximately 70% for TTS (text-to-speech), about 85% for Realtime voice services, and up to 95% for ASR (automatic speech recognition). The information was sourced from Jiemian News.
Read source