SAT, AUGUST 08, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ Large Language Models

Kimi K3 Review: 2.8 Trillion Parameters, $3/$15/M, and the DeepSeek Moment Round Two

Moonshot AI launched Kimi K3 at Shanghai World AI Conference July 16, 2026 — 2.8T params, $3/$15/M ($0.30/M cached), 1M context on Allegretto+. AA Index #4 of 189 globally (score 57). SWE Marathon 42.0% #1. Design Arena #1 frontend (1679 Elo). Open weights by July 27. Tiers: Moderato (256K, ~$20/mo), Allegretto+ (1M, ~$30-40/mo), Vivace (1M, top rate limits). China NI Law applies to hosted API.

By AIToolsRecap July 21, 2026 8 min read 1225 views
Home Articles Large Language Models Kimi Kimi K3 Review: 2.8 Trillion Parameters, $3/$15...

KIMI K3 — KEY FACTS

Parameters: 2.8 trillion total (896 experts, 16 active per token — Stable LatentMoE architecture)
Launched: July 16, 2026 at Shanghai World AI Conference
API pricing: $3/$15/M input/output ($0.30/M cached input)
Context: 1M tokens on Allegretto+ and Vivace tiers — 256K on Moderato
AA Intelligence Index: Score 57 — #4 of 189 models globally
SWE Marathon: 42.0% — #1 globally (vendor harness)
Design Arena: 1679 Elo — #1 globally for frontend coding
SWE-bench Pro: Not yet published
Open weights: By July 27, 2026
Data residency: Chinese company — China National Intelligence Law applies to hosted API

What Kimi K3 Actually Is

Kimi K3 is a 2.8 trillion parameter mixture-of-experts model built on Moonshot AI's Stable LatentMoE architecture. It uses 896 experts with 16 active per token — meaning at inference time, only 16 of the 896 expert networks are engaged for any given token. This design gives K3 the capability of a 2.8T model at the computational cost of roughly a 50B dense model per token, which is why the API pricing ($3/$15/M) is far below what a 2.8T dense model would cost to serve. It is the same architectural principle that DeepSeek used to make high-parameter models economically viable — applied at a larger scale than any previous open-weight release.

The model was announced at Shanghai's World AI Conference on July 16, 2026 and became available via API immediately. Open weights are confirmed for release by July 27 — seven days after announcement. Access points: kimi.com web interface, the Kimi iOS and Android apps, the Kimi Work desktop client, Kimi Code (the agentic coding product), and the Kimi API. Default thinking intensity is set to maximum on all access points.

Benchmarks — Where K3 Leads and Where It Has Gaps

Benchmark Kimi K3 Claude Fable 5 GPT-5.6 Sol Claude Opus 4.8
AA Intelligence Index 57 (#4) 60 (#1) 58.9 (#2) 56 (#5)
SWE Marathon 42.0% #1 24.0% N/A 26.0%
SWE-bench Pro Not published 80.4% #1 Not published 69.2%
Design Arena (frontend) #1 (1679 Elo) N/A N/A N/A
Input price /1M $3 ($0.30 cached) $10 $5 $5

SWE Marathon from Moonshot vendor harness — neutral replication pending. SWE-bench Pro for K3 not yet published. AA Index from Artificial Analysis Intelligence Index v4.1, July 2026.

Pricing and Subscription Tiers

Kimi K3 is accessible via API at $3/$15/M input/output. Cached input drops to $0.30/M — a 90% reduction for repeated context. This puts K3 40% cheaper than Claude Opus 4.8 ($5/$25/M) and 40% cheaper than GPT-5.6 Sol ($5/$30/M output) on a per-token basis.

The subscription tiers change how much context window you get with K3:

Moderato (~$20/month): Kimi K3 at 256K context. Entry tier. Good for standard conversational and coding tasks that fit within 256K tokens. Comparable to Claude Pro at the same price but with different model strengths.

Allegretto+ (~$30-40/month): Kimi K3 at full 1M context. The tier where K3's 1M context window advantage becomes relevant — large codebase ingestion, long documents, extended conversations. The right comparison tier vs Claude Max ($100/month) or SuperGrok Heavy ($60/month).

Vivace (top tier): Kimi K3 at 1M context with highest rate limits and priority queue access. For power users who consistently hit Allegretto rate limits. Same K3 model, higher throughput.

The DeepSeek Parallel — Why This Matters Beyond the Benchmarks

When DeepSeek launched in early 2025, it proved that a Chinese lab could ship a model matching GPT-4 performance at a fraction of the cost — and that the open-source community would adopt it at scale. K3 is attempting the next rung: beat the current best proprietary model (Claude Opus 4.8) on overall intelligence benchmarks, then release the weights so AWS, Azure, GCP, and a dozen smaller cloud providers can host competing versions. The open-weight release on July 27 is specifically designed to let Western-infrastructure hosting eliminate the data residency concern — a direct response to the criticism that hosted Chinese AI models carry regulatory risk under China's National Intelligence Law.

The 10-day gap between API launch (July 16) and weights release (July 27) is notable. It gives Moonshot AI 10 days of hosted API revenue and benchmark press before the model becomes a commodity that anyone can self-host. That sequencing is identical to what DeepSeek did with V3 and V4.

Data Residency — The Critical Consideration

Moonshot AI is a Chinese company. China's National Intelligence Law (Article 7, 2017) requires Chinese companies and citizens to cooperate with government intelligence requests on demand, with no public notification requirement. This applies to data processed by Moonshot's hosted Kimi API. For finance, healthcare, defence, legal, and any regulated workload involving sensitive IP or personal data: do not use the Kimi hosted API until legal review confirms it is acceptable for your jurisdiction.

After July 27, self-hosting K3 on AWS, Azure, or GCP eliminates this concern — you control where the model runs and what data touches it. For teams with GPU access, self-hosted K3 at Allegretto+ capability level with Western data residency changes the risk profile entirely. A 2.8T MoE model at Q4 quantisation requires approximately 200-400GB VRAM at minimum — plan GPU provisioning accordingly before the weights ship.

Who Should Use Kimi K3

Frontend developers: Design Arena #1 globally at 1679 Elo. K3 produces better frontend code than any other model in public evaluation. If you build UIs, React components, or landing pages with AI assistance, K3 is the clearest upgrade available today.

Long-horizon agentic coding teams: SWE Marathon 42.0% #1 — the 16-percentage-point lead over Fable 5 (24.0%) on long-horizon autonomous coding is the largest performance gap between K3 and Western frontier models. For multi-step agent tasks running over extended sessions, K3 has no published peer.

Cost-sensitive high-volume API users (non-regulated): At $3/$15/M ($0.30/M cached), K3 is 40-65% cheaper than comparable Western frontier models. For consumer-facing applications, research workloads, or any high-volume pipeline where data sovereignty is not a constraint, K3 offers the best price-per-intelligence available.

Not recommended for: Regulated industries (finance, healthcare, defence, legal), proprietary source code, customer PII, or any workload where Chinese data residency creates legal risk. Wait for July 27 weights and self-host on Western infrastructure instead.

Sources: Moonshot AI official blog · Artificial Analysis Intelligence Index v4.1 · Windows News · Latent Space (AINews July 2026) · Digital Applied · Related: Kimi K3 vs Claude Opus 4.8 → · Kimi K3 vs GPT-5.6 Sol → · Kimi Allegretto vs Claude Max → · Kimi Moderato vs Claude Pro →

Tags
AI NewsGenerative AICoding AI2026Best AI Tools

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →
💡 Kimi prompts
Prompt Guide
Best Kimi K3 Prompts for Coding (2026)
Kimi K3 ranks #1 globally on Design Arena frontend coding (1679 Elo) and #1 on SWE Marathon (42.0%) — making it the strongest available model for frontend development and long-horizon agentic coding tasks. At $3/$15/M with 1M context on Allegretto+ tier, it offers the best price-to-coding-capability ratio of any frontier model. These prompts are optimised for Kimi Code CLI, the Kimi API, and kimi.com on Allegretto+ or Vivace tiers.
Get Prompts →
Prompt Guide
Best Kimi K3 Prompts for Research (2026)
Kimi K3's 1M context window on Allegretto+ tier makes it one of the strongest research models available — it can ingest entire research papers, technical documentation sets, and large code repositories in a single session. At $0.30/M cached input, repeated context (like a standing research brief or knowledge base) becomes extremely cost-efficient. These prompts are optimised for kimi.com on Allegretto+ or Vivace tiers and the Kimi API.
Get Prompts →
Prompt Guide
Best Kimi K3 Prompts for Writing (2026)
Kimi K3 on Allegretto+ tier offers 1M context for writing tasks that require deep document awareness — editing a long manuscript, maintaining style consistency across a large content set, or rewriting with full context of everything written before. At $0.30/M cached input, repeatedly referencing a style guide or brand voice document costs almost nothing. These prompts are optimised for kimi.com on Allegretto+ and the Kimi API.
Get Prompts →