SAT, SEPTEMBER 26, 2026
Independent · In‑Depth · Practitioner‑Tested
Voice & Audio

ElevenLabs vs Qwen-Audio 3.1 (2026): What 95% Off Actually Buys

Alibaba cut voice API prices by up to 95% on 25 September. ElevenLabs sits at $100 per million characters. The gap is real - and so is the thing you give up.

🕐 7 min read 👁 13 views 📅 Sep 26, 2026

The Short Version

On 25 September 2026 Alibaba shipped Qwen-Audio 3.1: five models, with price cuts of up to 95% on speech recognition, about 70% on text-to-speech and roughly 85% on realtime interaction.

The headline comparison making the rounds is that Qwen TTS undercuts ElevenLabs by roughly 70%, measured against ElevenLabs' $100 per million characters.

That claim is directionally sound and it is not a quote you can budget from, because Alibaba published percentages and not per-unit prices, and stated no effective date or whether the discounted tier is capacity-committed or spot.

What Each One Is

ElevenLabs

A voice specialist. Its position is English-language naturalness, voice cloning, and a production track record in audiobooks, dubbing and media. $100 per million characters on the reference tier.

Qwen-Audio 3.1

A five-model stack, not a single product:

  • ASR - speech recognition, upgraded, cut up to 95%
  • TTS - text-to-speech, upgraded, cut about 70%
  • Realtime - live interaction, upgraded, cut roughly 85%
  • ASR-Next - new, for understanding audio rather than transcribing it
  • TTS-Next - new, for audio creation

TTS covers 16 languages and 20 Chinese dialect regions. Realtime Plus carries a 262K token context window.

Where the Comparison Is Strong

Chinese and Chinese-dialect audio

20 dialect regions is coverage no Western provider offers. If your audio is Mandarin, Cantonese or regional Chinese, this is not a price comparison - it is a capability comparison Qwen wins.

High-volume transcription

A 95% cut on ASR changes what is economically viable rather than just cheaper. Transcription also has the lowest switching cost of anything here, because the output is text and you can diff two providers on the same audio in an afternoon.

Audio understanding as a separate job

ASR-Next answering questions about audio directly, rather than transcribing then feeding a text model, removes a step and the transcription errors that step propagates. ElevenLabs does not position a product here.

Where It Is Weak

English production voice

ElevenLabs' lead is English naturalness and that is exactly what a 70% saving would be buying you out of. Run your own copy through both before moving anything customer-facing.

Language breadth

16 languages is narrower than ElevenLabs' coverage. If you publish in more than 16, the comparison ends there.

The price itself

"Up to 95%" is a marketing range, not a rate card. "Up to" is doing real work - the maximum cut applies to some configuration, not necessarily yours. Until Alibaba publishes per-unit pricing, get a quote for your actual volume and usage pattern.

What Is Genuinely New, Beyond Price

Qwen TTS is prompt-steerable for delivery: you instruct it with something like "read this with a sharp, commanding tone" rather than picking a fixed voice preset. The realtime model adjusts pace and tone to detected emotional state, slowing down when it reads a lower state.

Whether that beats a well-chosen ElevenLabs voice is an open question and a testable one. It is a different control surface, not obviously a better one.

Decision Framework

  • Chinese or dialect audio - Qwen-Audio 3.1. Not close.
  • Bulk transcription at volume - Qwen. Test on your own audio; switching costs are near zero.
  • English audiobooks, dubbing, brand voice - ElevenLabs. The saving is real and so is what you trade for it.
  • More than 16 publication languages - ElevenLabs.
  • Realtime voice agents - price the 262K context on Realtime Plus against your session lengths. That specification matters more than the 85% cut.
  • You need a quotable rate for a budget today - ElevenLabs. It publishes one.

The Market Context

This landed in the same week DeepSeek reported a $1 billion annualised run rate after raising API prices 2.3x to 4.5x. Two Chinese providers moved in opposite directions on price in seven days.

That is not a contradiction, it is two positions in one cycle. Alibaba is buying share in audio, where it does not have it. DeepSeek already has share in text and is testing what it is worth. The practical warning is the same in both directions: a price set to win a market is not a price that survives winning it.

Verdict

Qwen-Audio 3.1 is the clear choice for Chinese-language audio and for transcription at volume, where the cut is large enough to change what you can afford to build. ElevenLabs holds English production voice and language breadth, and it has the advantage of publishing a number you can put in a budget.

Do not migrate production English voice on a percentage. Wait for Alibaba's per-unit pricing, then test on your own copy.

⚖ Our Verdict

Qwen-Audio 3.1 wins Chinese-language audio outright and wins bulk transcription on price, with cuts of up to 95% on ASR. ElevenLabs keeps English production voice and broader language coverage, and publishes an actual rate card - Alibaba announced percentages, not per-unit prices, and gave no effective date. Test before migrating anything customer-facing.