KIMI K3 WEIGHTS — NOW LIVE (JULY 27, 2026)
● Location: huggingface.co/moonshotai/Kimi-K3
● License: Modified MIT — commercial use permitted
● Download size: ~594GB MXFP4 safetensors (native format)
● Serving: vLLM with KDA prefill cache support available at launch
● Technical report: Published alongside weights
● Architecture: 2.8T params, 896 experts, 16 active per token, Kimi Delta Attention + Attention Residuals
● Minimum load: 8× H100 80GB (experimental) — 64+ accelerators for production
● Community quants: GGUF Q4 builds expected within 24h on r/LocalLLaMA
● Managed inference: Check Together AI, Fireworks AI, Groq — likely live within 48h
● Hallucination warning: Independent testing found 51% rate not in Moonshot benchmarks
The Modified MIT License — What You Can Do
Kimi K3 open weights arrive July 27 under a Modified MIT license. The Modified MIT license allows commercial use, self-hosting, fine-tuning, and redistribution. Read the specific terms on the model card before building a product — Modified MIT can include specific restrictions or attribution requirements that differ from standard MIT. For most enterprise deployment scenarios — self-hosting on AWS, Azure, or GCP Western regions, building products on top of K3, fine-tuning for specific domains — Modified MIT is permissive enough to proceed. Verify the exact license file in the repository for edge cases such as model distillation, derivative model distribution, and white-labelling.
K3 is the next generation of open frontier models from Moonshot AI. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning. The architecture is built on Kimi Delta Attention and Attention Residuals, with native agentic capabilities including tool calling, browsing, and multi-step planning. The technical report published alongside the weights covers architecture details, training methodology, and evaluation results — the first complete public documentation of K3's internals.
Download and Setup — Step by Step
☐ Step 1: Confirm the repo is under the verified Moonshot organization at huggingface.co/moonshotai/Kimi-K3 — not a copycat
☐ Step 2: Read the LICENSE file before commercial deployment — confirm Modified MIT terms for your use case
☐ Step 3: huggingface-cli download moonshotai/Kimi-K3 — ensure 1.2TB+ free storage (download + inference cache)
☐ Step 4: pip install -U vllm — KDA prefill cache support ships with weights in latest vLLM
☐ Step 5 (if on H100, not Blackwell): Check the model card for MXFP4 weight conversion steps — H100 may require a conversion pass before loading
☐ Step 6: If you do not have GPU infrastructure — check Together AI, Fireworks AI, and Groq model pages. Managed K3 inference on Western infrastructure eliminates the China NI Law concern without self-hosting overhead.
The Hallucination Finding — Do Not Skip This
Independent testing found a 51% hallucination rate that Moonshot AI omitted from its benchmark charts. This is not in the official K3 benchmark materials. For workloads where factual accuracy is critical — RAG pipelines, customer support, knowledge base queries, legal or financial applications — run your own hallucination benchmark before production deployment. K3's strengths are real and verified: SWE Marathon #1 (agentic coding), Design Arena #1 (frontend). The hallucination risk is also real and independently documented. Both are true simultaneously.
Why This Release Is Historically Significant
By mid 2026, Chinese AI labs are in frontier models. It is becoming clear that the Chinese labs are far more capital efficient. In a world where scaling laws dictate that intelligence is proportional to effective capital, that may be the greatest strength your AI industry could ever have. K3 at 2.8T parameters is the largest open-weight model ever released by a factor of approximately 2× — the previous record was DeepSeek V4 at approximately 1.6T. K3 delivers near-frontier capability at roughly one-third the cost, plus two things closed models cannot offer: downloadable weights and a 1 million token context window. A team that self-hosts K3 on Western cloud infrastructure has access to a model that matches Claude Opus 4.8 on most benchmarks, costs approximately $0.14/1K tokens at Q4 on reserved H100s, and is under no vendor control — no Musk clause, no export control risk, no subscription changes.
What the White House Accusation Means (Do Not Ignore This)
The White House has publicly accused Moonshot of using thousands of fraudulent accounts to harvest Claude conversations for training data. This accusation — if substantiated — means K3's training data may include Claude outputs obtained without Anthropic's permission. For enterprise legal teams evaluating K3 deployment: this accusation is not resolved, Moonshot has not publicly responded to it, and it is relevant to any vendor risk assessment. The Modified MIT license covers your right to use the weights — it does not resolve questions about how the training data was obtained.
Sources: Hugging Face moonshotai/Kimi-K3 · TechTimes · ExplainX · Wan27 · Interconnects.ai · TECHi · NodeMini · Related: Full self-hosting cost guide → · Kimi K3 vs Claude Opus 5 → · Kimi K3 full review →