MON, JULY 27, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ Large Language Models

Kimi K3 Weights Are Live — Download From HuggingFace, Modified MIT License Confirmed, 51% Hallucination Warning

Kimi K3 open weights are live at huggingface.co/moonshotai/Kimi-K3 under Modified MIT license — commercial use permitted. ~594GB MXFP4 download. vLLM with KDA prefill cache available. Minimum: 8× H100 80GB to load. Community GGUF Q4 builds expected within 24h. Key warning: independent testing found 51% hallucination rate not disclosed in Moonshot's benchmark charts. White House accused Moonshot of using fraudulent accounts to harvest Claude training data — not yet resolved.

By AIToolsRecap July 27, 2026 6 min read 64 views
Home Articles Large Language Models Kimi Kimi K3 Weights Are Live: Download From Hugging...

KIMI K3 WEIGHTS — NOW LIVE (JULY 27, 2026)

Location: huggingface.co/moonshotai/Kimi-K3
License: Modified MIT — commercial use permitted
Download size: ~594GB MXFP4 safetensors (native format)
Serving: vLLM with KDA prefill cache support available at launch
Technical report: Published alongside weights
Architecture: 2.8T params, 896 experts, 16 active per token, Kimi Delta Attention + Attention Residuals
Minimum load: 8× H100 80GB (experimental) — 64+ accelerators for production
Community quants: GGUF Q4 builds expected within 24h on r/LocalLLaMA
Managed inference: Check Together AI, Fireworks AI, Groq — likely live within 48h
Hallucination warning: Independent testing found 51% rate not in Moonshot benchmarks

The Modified MIT License — What You Can Do

Kimi K3 open weights arrive July 27 under a Modified MIT license. The Modified MIT license allows commercial use, self-hosting, fine-tuning, and redistribution. Read the specific terms on the model card before building a product — Modified MIT can include specific restrictions or attribution requirements that differ from standard MIT. For most enterprise deployment scenarios — self-hosting on AWS, Azure, or GCP Western regions, building products on top of K3, fine-tuning for specific domains — Modified MIT is permissive enough to proceed. Verify the exact license file in the repository for edge cases such as model distillation, derivative model distribution, and white-labelling.

K3 is the next generation of open frontier models from Moonshot AI. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning. The architecture is built on Kimi Delta Attention and Attention Residuals, with native agentic capabilities including tool calling, browsing, and multi-step planning. The technical report published alongside the weights covers architecture details, training methodology, and evaluation results — the first complete public documentation of K3's internals.

Download and Setup — Step by Step

Step 1: Confirm the repo is under the verified Moonshot organization at huggingface.co/moonshotai/Kimi-K3 — not a copycat

Step 2: Read the LICENSE file before commercial deployment — confirm Modified MIT terms for your use case

Step 3: huggingface-cli download moonshotai/Kimi-K3 — ensure 1.2TB+ free storage (download + inference cache)

Step 4: pip install -U vllm — KDA prefill cache support ships with weights in latest vLLM

Step 5 (if on H100, not Blackwell): Check the model card for MXFP4 weight conversion steps — H100 may require a conversion pass before loading

Step 6: If you do not have GPU infrastructure — check Together AI, Fireworks AI, and Groq model pages. Managed K3 inference on Western infrastructure eliminates the China NI Law concern without self-hosting overhead.

The Hallucination Finding — Do Not Skip This

Independent testing found a 51% hallucination rate that Moonshot AI omitted from its benchmark charts. This is not in the official K3 benchmark materials. For workloads where factual accuracy is critical — RAG pipelines, customer support, knowledge base queries, legal or financial applications — run your own hallucination benchmark before production deployment. K3's strengths are real and verified: SWE Marathon #1 (agentic coding), Design Arena #1 (frontend). The hallucination risk is also real and independently documented. Both are true simultaneously.

Why This Release Is Historically Significant

By mid 2026, Chinese AI labs are in frontier models. It is becoming clear that the Chinese labs are far more capital efficient. In a world where scaling laws dictate that intelligence is proportional to effective capital, that may be the greatest strength your AI industry could ever have. K3 at 2.8T parameters is the largest open-weight model ever released by a factor of approximately 2× — the previous record was DeepSeek V4 at approximately 1.6T. K3 delivers near-frontier capability at roughly one-third the cost, plus two things closed models cannot offer: downloadable weights and a 1 million token context window. A team that self-hosts K3 on Western cloud infrastructure has access to a model that matches Claude Opus 4.8 on most benchmarks, costs approximately $0.14/1K tokens at Q4 on reserved H100s, and is under no vendor control — no Musk clause, no export control risk, no subscription changes.

What the White House Accusation Means (Do Not Ignore This)

The White House has publicly accused Moonshot of using thousands of fraudulent accounts to harvest Claude conversations for training data. This accusation — if substantiated — means K3's training data may include Claude outputs obtained without Anthropic's permission. For enterprise legal teams evaluating K3 deployment: this accusation is not resolved, Moonshot has not publicly responded to it, and it is relevant to any vendor risk assessment. The Modified MIT license covers your right to use the weights — it does not resolve questions about how the training data was obtained.

Sources: Hugging Face moonshotai/Kimi-K3 · TechTimes · ExplainX · Wan27 · Interconnects.ai · TECHi · NodeMini · Related: Full self-hosting cost guide → · Kimi K3 vs Claude Opus 5 → · Kimi K3 full review →

Tags
KimiAI NewsGenerative AICoding AI2026

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →
💡 Kimi prompts
Prompt Guide
Best Kimi K3 Prompts for Coding (2026)
Kimi K3 ranks #1 globally on Design Arena frontend coding (1679 Elo) and #1 on SWE Marathon (42.0%) — making it the strongest available model for frontend development and long-horizon agentic coding tasks. At $3/$15/M with 1M context on Allegretto+ tier, it offers the best price-to-coding-capability ratio of any frontier model. These prompts are optimised for Kimi Code CLI, the Kimi API, and kimi.com on Allegretto+ or Vivace tiers.
Get Prompts →
Prompt Guide
Best Kimi K3 Prompts for Research (2026)
Kimi K3's 1M context window on Allegretto+ tier makes it one of the strongest research models available — it can ingest entire research papers, technical documentation sets, and large code repositories in a single session. At $0.30/M cached input, repeated context (like a standing research brief or knowledge base) becomes extremely cost-efficient. These prompts are optimised for kimi.com on Allegretto+ or Vivace tiers and the Kimi API.
Get Prompts →
Prompt Guide
Best Kimi K3 Prompts for Writing (2026)
Kimi K3 on Allegretto+ tier offers 1M context for writing tasks that require deep document awareness — editing a long manuscript, maintaining style consistency across a large content set, or rewriting with full context of everything written before. At $0.30/M cached input, repeatedly referencing a style guide or brand voice document costs almost nothing. These prompts are optimised for kimi.com on Allegretto+ and the Kimi API.
Get Prompts →