What shipped
Alibaba released Qwen-Audio 3.1 on 25 September 2026 - a five-model audio stack with price cuts across the lineup.
- ASR - speech recognition, upgraded. Price cut up to 95%.
- TTS - text-to-speech, upgraded. Price cut about 70%.
- Realtime - live interaction, upgraded. Price cut roughly 85%.
- ASR-Next - new, positioned for audio understanding rather than transcription.
- TTS-Next - new, positioned for audio creation.
TTS covers 16 languages and 20 Chinese dialect regions. Realtime Plus carries a 262K token context window.
The number Alibaba did not publish
Alibaba announced percentage reductions, not per-unit prices, and did not state an effective date or say whether the discounted tier is capacity-committed or spot.
That matters when you are comparing. "Up to 95% off" is not a price, and "up to" is doing real work in that sentence - the cut applies to some configuration, not necessarily the one you would use. Before migrating, get the quote for your actual usage pattern rather than the headline reduction.
The one comparison available is on TTS, which is reported to undercut ElevenLabs by roughly 70% against ElevenLabs' $100 per million characters.
What is actually new, beyond price
TTS is prompt-steerable for delivery: you can instruct it with something like "read this with a sharp, commanding tone" rather than selecting a fixed voice preset. The realtime model adjusts pace and tone based on detected emotional state, responding more slowly when it reads a lower state.
Splitting ASR into transcription and understanding is the more interesting structural change. ASR-Next is aimed at answering questions about audio rather than converting it to text, which is a different job and has usually meant piping a transcript into a text model.
Should you switch?
- High-volume transcription - worth pricing. A 95% cut on ASR is large enough that it changes what is economically viable, and transcription has low switching costs because output is easy to compare.
- Production voice in English - test before moving. ElevenLabs' quality lead in English is the thing you would be trading away, and 16 languages is narrower than ElevenLabs' coverage.
- Chinese or Chinese-dialect audio - strongest case. 20 dialect regions is coverage no Western provider matches.
- Realtime voice agents - the 262K context on Realtime Plus is the specification to check against your use case, more than the 85% cut.
Context
This landed the same week DeepSeek reported a $1 billion run rate after raising prices 2.3x to 4.5x. Two Chinese providers moved in opposite directions on price in the same week, which is what a maturing market looks like: DeepSeek is monetising a position it already holds, Alibaba is buying one it does not.
FAQ
How much does Qwen-Audio 3.1 cost?
Alibaba published percentage cuts - up to 95% on ASR, about 70% on TTS, roughly 85% on realtime - but not per-unit prices. Check Qwen Cloud for current rates.
How many models are in the stack?
Five: ASR, TTS and Realtime upgraded, plus new ASR-Next and TTS-Next.
How does TTS compare to ElevenLabs on price?
Roughly 70% cheaper, against ElevenLabs at $100 per million characters.
What languages does it support?
16 languages for TTS, plus 20 Chinese dialect regions.