WED, SEPTEMBER 30, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ General

China's CUDA Alternative Shipped Without a Single Benchmark

DeepSeek and Huawei released TileLang, DeepGEMM Ascend and DeepEP Ascend for Ascend 950 accelerators - the layer that actually holds developers on Nvidia - and published no performance comparison at all.

By AIToolsRecap September 30, 2026 6 min read 11 views
Home › Articles › General › DeepSeek Open-Sourced a CUDA Alternative for Hu...

What Was Released

On 30 September 2026, DeepSeek open-sourced programming infrastructure for Huawei's Ascend platform, developed with Huawei. Three components:

  • TileLang - a programming language positioned as an alternative to Nvidia's CUDA, described as offering a simpler programming model
  • DeepGEMM Ascend - a matrix multiplication kernel library for high-throughput operations
  • DeepEP Ascend - efficient large-scale communication across multiple devices

The tools target Ascend 950 chips. DeepSeek is separately planning to deploy over 160,000 Huawei Ascend chips at a new data centre in Inner Mongolia.

Why This Is About Software, Not Chips

Nvidia's position has never rested mainly on having the fastest silicon. It rests on CUDA - roughly two decades of libraries, kernels, tooling and accumulated engineer familiarity. Competing chips have existed for years. Competing software ecosystems have not.

That is what makes this release more significant than another accelerator announcement. TileLang is an attempt at the layer that actually holds people in place, and DeepGEMM and DeepEP are the two pieces you need immediately after a language: fast matrix multiplication and multi-device communication. Between them, that is most of what training a large model requires from the stack.

Whether it works is a different question, and nobody can answer it yet.

No Benchmarks. At All.

DeepSeek says TileLang enables reaching the hardware's full performance potential. It has published no comparative figures against CUDA - no throughput numbers, no training time, no utilisation percentages.

For a release whose entire claim is that it is a viable CUDA alternative, that absence is the most important thing about it today. A simpler programming model that runs slower is a research project. A simpler programming model that matches CUDA on Ascend hardware changes the economics of Chinese AI training.

Until someone publishes numbers, treat this as a serious statement of intent rather than a demonstrated result. That is not scepticism about the engineering - DeepSeek has shipped genuinely strong work before. It is just what the available evidence supports.

The 160,000 Chips Are the Commitment

The Inner Mongolia deployment is the part that is not a claim. Ordering more than 160,000 Ascend accelerators is a capital decision that assumes this software stack will work, made by the company writing it.

Firms adopt their own tools all the time. But at that scale, the toolkit is not a side project - it is infrastructure DeepSeek itself depends on, which is a stronger signal than any announcement.

Two Attacks on the Same Moat, One Day Apart

This landed alongside AMD paying $8.2 billion for World Labs, partly to get a frontier model team developing on its ROCm stack.

Two very different companies, same week, same target - and neither is trying to out-engineer Nvidia's hardware. AMD is buying ecosystem credibility. DeepSeek and Huawei are building it in the open and giving it away.

The open-source route is slower and has a better track record.

What It Means For You

If you train outside China on Nvidia hardware, nothing changes this year. CUDA's advantage is cumulative and a first release does not dent it.

What is worth watching is whether anyone outside DeepSeek adopts TileLang, and whether independent benchmarks appear. An ecosystem is defined by its second and third users, not its first.

FAQ

What is TileLang?

A programming language released by DeepSeek for Huawei Ascend accelerators, positioned as a simpler alternative to Nvidia's CUDA.

What else was released?

DeepGEMM Ascend, a matrix multiplication kernel library, and DeepEP Ascend, for large-scale multi-device communication.

Which chips does it support?

Huawei Ascend, with optimisation specifically noted for the Ascend 950.

Is it faster than CUDA?

Unknown. DeepSeek has published no comparative benchmarks, only a claim that the tools allow reaching the hardware's full performance potential.

Is it free?

The tools have been released openly. A specific licence has not been confirmed in reporting.

Tags
AI NewsNvidiaCoding AI2026
⚑

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →
💡 AI Tools prompts
Prompt Guide
Best Claude AI Prompts for SEO (2026) — Content, Technical, and Comparison SEO
Claude Sonnet 5 and Opus 5 are strong for SEO work that requires writing quality, structured analysis, and long-form content generation. With 1M context, Claude can analyse an entire site's content structure, compare competing pages, and write complete article drafts in one session. These prompts cover the full SEO workflow: keyword research synthesis, content briefs, on-page optimisation, meta descriptions, technical audit interpretation, and comparison content that ranks above AI Overviews.
Get Prompts →
Prompt Guide
Best ChatGPT Prompts for SEO (2026) — GPT-5.6 and Browse
ChatGPT with GPT-5.6 Sol and Browse enabled is a capable SEO research tool — it can search the live web, analyse SERP results, and synthesise content briefs in a single session. GPT-5.6 Terra at $2.50/M offers a cost-efficient option for high-volume SEO content generation. These prompts are optimised for ChatGPT Plus with Browse, the ChatGPT Work product for larger projects, and the OpenAI API with web_search tool enabled.
Get Prompts →
Prompt Guide
Best Claude Opus 5 and Sonnet 5 Prompts for Writing (2026)
Claude Opus 5 and Sonnet 5 consistently produce the highest-quality long-form writing of any AI model in July 2026 — a lead documented across writing benchmarks and user testing since Claude 3 Opus. With 1M context and 128K output on Opus 5, Claude can write book chapters, complete reports, and long-form content without truncating. Sonnet 5 at $2/$10/M (intro through August 31) is the best value writing model available. These prompts are optimised for claude.ai Pro/Max, Claude Cowork, and the API.
Get Prompts →