KIMI K3 — KEY FACTS
● Parameters: 2.8 trillion total (896 experts, 16 active per token — Stable LatentMoE architecture)
● Launched: July 16, 2026 at Shanghai World AI Conference
● API pricing: $3/$15/M input/output ($0.30/M cached input)
● Context: 1M tokens on Allegretto+ and Vivace tiers — 256K on Moderato
● AA Intelligence Index: Score 57 — #4 of 189 models globally
● SWE Marathon: 42.0% — #1 globally (vendor harness)
● Design Arena: 1679 Elo — #1 globally for frontend coding
● SWE-bench Pro: Not yet published
● Open weights: By July 27, 2026
● Data residency: Chinese company — China National Intelligence Law applies to hosted API
What Kimi K3 Actually Is
Kimi K3 is a 2.8 trillion parameter mixture-of-experts model built on Moonshot AI's Stable LatentMoE architecture. It uses 896 experts with 16 active per token — meaning at inference time, only 16 of the 896 expert networks are engaged for any given token. This design gives K3 the capability of a 2.8T model at the computational cost of roughly a 50B dense model per token, which is why the API pricing ($3/$15/M) is far below what a 2.8T dense model would cost to serve. It is the same architectural principle that DeepSeek used to make high-parameter models economically viable — applied at a larger scale than any previous open-weight release.
The model was announced at Shanghai's World AI Conference on July 16, 2026 and became available via API immediately. Open weights are confirmed for release by July 27 — seven days after announcement. Access points: kimi.com web interface, the Kimi iOS and Android apps, the Kimi Work desktop client, Kimi Code (the agentic coding product), and the Kimi API. Default thinking intensity is set to maximum on all access points.
Benchmarks — Where K3 Leads and Where It Has Gaps
| Benchmark |
Kimi K3 |
Claude Fable 5 |
GPT-5.6 Sol |
Claude Opus 4.8 |
| AA Intelligence Index |
57 (#4) |
60 (#1) |
58.9 (#2) |
56 (#5) |
| SWE Marathon |
42.0% #1 |
24.0% |
N/A |
26.0% |
| SWE-bench Pro |
Not published |
80.4% #1 |
Not published |
69.2% |
| Design Arena (frontend) |
#1 (1679 Elo) |
N/A |
N/A |
N/A |
| Input price /1M |
$3 ($0.30 cached) |
$10 |
$5 |
$5 |
SWE Marathon from Moonshot vendor harness — neutral replication pending. SWE-bench Pro for K3 not yet published. AA Index from Artificial Analysis Intelligence Index v4.1, July 2026.
Pricing and Subscription Tiers
Kimi K3 is accessible via API at $3/$15/M input/output. Cached input drops to $0.30/M — a 90% reduction for repeated context. This puts K3 40% cheaper than Claude Opus 4.8 ($5/$25/M) and 40% cheaper than GPT-5.6 Sol ($5/$30/M output) on a per-token basis.
The subscription tiers change how much context window you get with K3:
Moderato (~$20/month): Kimi K3 at 256K context. Entry tier. Good for standard conversational and coding tasks that fit within 256K tokens. Comparable to Claude Pro at the same price but with different model strengths.
Allegretto+ (~$30-40/month): Kimi K3 at full 1M context. The tier where K3's 1M context window advantage becomes relevant — large codebase ingestion, long documents, extended conversations. The right comparison tier vs Claude Max ($100/month) or SuperGrok Heavy ($60/month).
Vivace (top tier): Kimi K3 at 1M context with highest rate limits and priority queue access. For power users who consistently hit Allegretto rate limits. Same K3 model, higher throughput.
The DeepSeek Parallel — Why This Matters Beyond the Benchmarks
When DeepSeek launched in early 2025, it proved that a Chinese lab could ship a model matching GPT-4 performance at a fraction of the cost — and that the open-source community would adopt it at scale. K3 is attempting the next rung: beat the current best proprietary model (Claude Opus 4.8) on overall intelligence benchmarks, then release the weights so AWS, Azure, GCP, and a dozen smaller cloud providers can host competing versions. The open-weight release on July 27 is specifically designed to let Western-infrastructure hosting eliminate the data residency concern — a direct response to the criticism that hosted Chinese AI models carry regulatory risk under China's National Intelligence Law.
The 10-day gap between API launch (July 16) and weights release (July 27) is notable. It gives Moonshot AI 10 days of hosted API revenue and benchmark press before the model becomes a commodity that anyone can self-host. That sequencing is identical to what DeepSeek did with V3 and V4.
Data Residency — The Critical Consideration
Moonshot AI is a Chinese company. China's National Intelligence Law (Article 7, 2017) requires Chinese companies and citizens to cooperate with government intelligence requests on demand, with no public notification requirement. This applies to data processed by Moonshot's hosted Kimi API. For finance, healthcare, defence, legal, and any regulated workload involving sensitive IP or personal data: do not use the Kimi hosted API until legal review confirms it is acceptable for your jurisdiction.
After July 27, self-hosting K3 on AWS, Azure, or GCP eliminates this concern — you control where the model runs and what data touches it. For teams with GPU access, self-hosted K3 at Allegretto+ capability level with Western data residency changes the risk profile entirely. A 2.8T MoE model at Q4 quantisation requires approximately 200-400GB VRAM at minimum — plan GPU provisioning accordingly before the weights ship.
Who Should Use Kimi K3
Frontend developers: Design Arena #1 globally at 1679 Elo. K3 produces better frontend code than any other model in public evaluation. If you build UIs, React components, or landing pages with AI assistance, K3 is the clearest upgrade available today.
Long-horizon agentic coding teams: SWE Marathon 42.0% #1 — the 16-percentage-point lead over Fable 5 (24.0%) on long-horizon autonomous coding is the largest performance gap between K3 and Western frontier models. For multi-step agent tasks running over extended sessions, K3 has no published peer.
Cost-sensitive high-volume API users (non-regulated): At $3/$15/M ($0.30/M cached), K3 is 40-65% cheaper than comparable Western frontier models. For consumer-facing applications, research workloads, or any high-volume pipeline where data sovereignty is not a constraint, K3 offers the best price-per-intelligence available.
Not recommended for: Regulated industries (finance, healthcare, defence, legal), proprietary source code, customer PII, or any workload where Chinese data residency creates legal risk. Wait for July 27 weights and self-host on Western infrastructure instead.
Sources: Moonshot AI official blog · Artificial Analysis Intelligence Index v4.1 · Windows News · Latent Space (AINews July 2026) · Digital Applied · Related: Kimi K3 vs Claude Opus 4.8 → · Kimi K3 vs GPT-5.6 Sol → · Kimi Allegretto vs Claude Max → · Kimi Moderato vs Claude Pro →