SAT, SEPTEMBER 26, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ General

Alibaba Just Made Voice AI Cheap Enough to Reconsider

Qwen-Audio 3.1 ships five models - ASR, TTS, Realtime plus new ASR-Next and TTS-Next - with cuts of up to 95% on speech recognition, about 70% on TTS and roughly 85% on realtime, though Alibaba published percentages rather than per-unit prices.

By AIToolsRecap September 26, 2026 6 min read 14 views
Home › Articles › General › Alibaba Cut Voice API Prices by Up to 95%

What shipped

Alibaba released Qwen-Audio 3.1 on 25 September 2026 - a five-model audio stack with price cuts across the lineup.

  • ASR - speech recognition, upgraded. Price cut up to 95%.
  • TTS - text-to-speech, upgraded. Price cut about 70%.
  • Realtime - live interaction, upgraded. Price cut roughly 85%.
  • ASR-Next - new, positioned for audio understanding rather than transcription.
  • TTS-Next - new, positioned for audio creation.

TTS covers 16 languages and 20 Chinese dialect regions. Realtime Plus carries a 262K token context window.

The number Alibaba did not publish

Alibaba announced percentage reductions, not per-unit prices, and did not state an effective date or say whether the discounted tier is capacity-committed or spot.

That matters when you are comparing. "Up to 95% off" is not a price, and "up to" is doing real work in that sentence - the cut applies to some configuration, not necessarily the one you would use. Before migrating, get the quote for your actual usage pattern rather than the headline reduction.

The one comparison available is on TTS, which is reported to undercut ElevenLabs by roughly 70% against ElevenLabs' $100 per million characters.

What is actually new, beyond price

TTS is prompt-steerable for delivery: you can instruct it with something like "read this with a sharp, commanding tone" rather than selecting a fixed voice preset. The realtime model adjusts pace and tone based on detected emotional state, responding more slowly when it reads a lower state.

Splitting ASR into transcription and understanding is the more interesting structural change. ASR-Next is aimed at answering questions about audio rather than converting it to text, which is a different job and has usually meant piping a transcript into a text model.

Should you switch?

  • High-volume transcription - worth pricing. A 95% cut on ASR is large enough that it changes what is economically viable, and transcription has low switching costs because output is easy to compare.
  • Production voice in English - test before moving. ElevenLabs' quality lead in English is the thing you would be trading away, and 16 languages is narrower than ElevenLabs' coverage.
  • Chinese or Chinese-dialect audio - strongest case. 20 dialect regions is coverage no Western provider matches.
  • Realtime voice agents - the 262K context on Realtime Plus is the specification to check against your use case, more than the 85% cut.

Context

This landed the same week DeepSeek reported a $1 billion run rate after raising prices 2.3x to 4.5x. Two Chinese providers moved in opposite directions on price in the same week, which is what a maturing market looks like: DeepSeek is monetising a position it already holds, Alibaba is buying one it does not.

FAQ

How much does Qwen-Audio 3.1 cost?

Alibaba published percentage cuts - up to 95% on ASR, about 70% on TTS, roughly 85% on realtime - but not per-unit prices. Check Qwen Cloud for current rates.

How many models are in the stack?

Five: ASR, TTS and Realtime upgraded, plus new ASR-Next and TTS-Next.

How does TTS compare to ElevenLabs on price?

Roughly 70% cheaper, against ElevenLabs at $100 per million characters.

What languages does it support?

16 languages for TTS, plus 20 Chinese dialect regions.

Tags
AI NewsVoice AIGenerative AI2026
⚑

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →
💡 AI Tools prompts
Prompt Guide
Best Claude AI Prompts for SEO (2026) — Content, Technical, and Comparison SEO
Claude Sonnet 5 and Opus 5 are strong for SEO work that requires writing quality, structured analysis, and long-form content generation. With 1M context, Claude can analyse an entire site's content structure, compare competing pages, and write complete article drafts in one session. These prompts cover the full SEO workflow: keyword research synthesis, content briefs, on-page optimisation, meta descriptions, technical audit interpretation, and comparison content that ranks above AI Overviews.
Get Prompts →
Prompt Guide
Best ChatGPT Prompts for SEO (2026) — GPT-5.6 and Browse
ChatGPT with GPT-5.6 Sol and Browse enabled is a capable SEO research tool — it can search the live web, analyse SERP results, and synthesise content briefs in a single session. GPT-5.6 Terra at $2.50/M offers a cost-efficient option for high-volume SEO content generation. These prompts are optimised for ChatGPT Plus with Browse, the ChatGPT Work product for larger projects, and the OpenAI API with web_search tool enabled.
Get Prompts →
Prompt Guide
Best Claude Opus 5 and Sonnet 5 Prompts for Writing (2026)
Claude Opus 5 and Sonnet 5 consistently produce the highest-quality long-form writing of any AI model in July 2026 — a lead documented across writing benchmarks and user testing since Claude 3 Opus. With 1M context and 128K output on Opus 5, Claude can write book chapters, complete reports, and long-form content without truncating. Sonnet 5 at $2/$10/M (intro through August 31) is the best value writing model available. These prompts are optimised for claude.ai Pro/Max, Claude Cowork, and the API.
Get Prompts →