KIMI K3 WEIGHTS — DROPPING TONIGHT
● Time: July 27, 2026 00:00 UTC = July 26, 8:00 PM ET = July 27, 7:00 AM VN time
● Where: huggingface.co/moonshotai — watch for the Kimi-K3 or Kimi-K3-Instruct repo
● Download size: ~594GB (native MXFP4 safetensors)
● License: Modified MIT (expected, matching K2 series — confirm on model card)
● vLLM: KDA prefill cache support ships with weights
● Technical report: Published simultaneously with weights
● Minimum hardware: 8× H100 80GB to load (experimental). 64+ accelerators recommended for production.
● Community quants: GGUF Q4 and GGUF Q2 builds expected within 24 hours on r/LocalLLaMA
● Hallucination note: Independent testing found 51% hallucination rate not disclosed in Moonshot's benchmark charts
What to Do Right Now (Before 8PM ET)
☐ Verify the Moonshot HF org: huggingface.co/moonshotai — confirm the K3 repo appears under the verified Moonshot organization, not a copycat
☐ Read the license first: Do not assume Modified MIT — read the LICENSE file in the repo before deploying commercially. K2 was Modified MIT; K3 confirmation comes on release.
☐ Prepare storage: 594GB download + inference cache = plan for 1.2TB+ free storage minimum
☐ vLLM readiness: Update vLLM to latest — KDA attention prefill cache support ships with the weights. Run pip install -U vllm before the drop.
☐ If waiting for quants: Watch r/LocalLLaMA and Bartowski's HF profile — community GGUF Q4 builds typically appear within 12-24 hours of a major weight release
☐ Check managed inference: Together AI, Fireworks AI, Groq — monitor their model pages from 8PM ET. Managed K3 inference on Western infrastructure eliminates the China NI Law concern without self-hosting overhead.
The Hallucination Finding — Read Before Deploying
Independent testing found a 51% hallucination rate that Moonshot AI omitted from its benchmark charts. This is a significant finding that is not disclosed in Moonshot's official K3 benchmark materials. The hallucination rate is particularly relevant for retrieval-augmented generation (RAG) pipelines, customer-facing applications, and any workload where factual accuracy is a hard requirement. K3's strengths — SWE Marathon #1 on agentic coding, Design Arena #1 on frontend coding — are real and independently verified. But the hallucination rate means K3 should be tested on factual accuracy benchmarks relevant to your use case before production deployment.
Hardware Reality — What Can Actually Run K3
| Setup | Can run K3? | Notes |
| RTX 4090 / Mac Studio | No | Cannot load even Q4 — 1.4TB VRAM minimum for Q4 |
| 8× H100 80GB | Experimental | Minimum load, very limited batch/context |
| 18-24× H100 80GB (Q4) | Production Q4 | ~$50/hr reserved — practical deployment |
| 64+ accelerators (full) | Full quality | Moonshot recommended minimum for production |
Why This Release Is Different From Other Open-Weight Drops
At 2.8 trillion parameters, Kimi K3 is the world's first 3T-class open-weights model. DeepSeek V4 Pro is approximately 1.6T parameters. Llama 3.1 405B is 405B. The scale difference is significant: K3 at full precision requires infrastructure that no previous open-weight model has required. The MXFP4 quantisation-aware training means Q4 quality degradation should be lower than models not trained this way — but the VRAM requirements remain extreme compared to any previous open-weight release. This is not a "run it on your gaming PC" release. It is a server-infrastructure release that happens to have open weights.
Sources: TechTimes · Wan27 · TechTimes (hallucination) · ExplainX · NodeMini · ByteIota · Related: Full self-hosting guide with cost estimates → · Kimi K3 full review → · Kimi K3 vs GPT-5.6 Sol →