SUN, JULY 26, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ Large Language Models

Kimi K3 Weights Drop Tonight 8PM ET — 594GB Download, Modified MIT License, 51% Hallucination Warning

Kimi K3 open weights go live tonight July 26 at 8PM ET (July 27 00:00 UTC) on huggingface.co/moonshotai. Download: ~594GB MXFP4 safetensors. License: Modified MIT (confirm on model card before commercial deployment). vLLM with KDA prefill cache ships with weights. Minimum hardware: 8× H100 80GB. Key warning: independent testing found a 51% hallucination rate not disclosed in Moonshot's official benchmark charts — test on factual tasks before production deployment.

By AIToolsRecap July 26, 2026 5 min read 47 views
Home Articles Large Language Models Kimi Kimi K3 Weights Drop Tonight at 8PM ET (July 27...

KIMI K3 WEIGHTS — DROPPING TONIGHT

Time: July 27, 2026 00:00 UTC = July 26, 8:00 PM ET = July 27, 7:00 AM VN time
Where: huggingface.co/moonshotai — watch for the Kimi-K3 or Kimi-K3-Instruct repo
Download size: ~594GB (native MXFP4 safetensors)
License: Modified MIT (expected, matching K2 series — confirm on model card)
vLLM: KDA prefill cache support ships with weights
Technical report: Published simultaneously with weights
Minimum hardware: 8× H100 80GB to load (experimental). 64+ accelerators recommended for production.
Community quants: GGUF Q4 and GGUF Q2 builds expected within 24 hours on r/LocalLLaMA
Hallucination note: Independent testing found 51% hallucination rate not disclosed in Moonshot's benchmark charts

What to Do Right Now (Before 8PM ET)

Verify the Moonshot HF org: huggingface.co/moonshotai — confirm the K3 repo appears under the verified Moonshot organization, not a copycat

Read the license first: Do not assume Modified MIT — read the LICENSE file in the repo before deploying commercially. K2 was Modified MIT; K3 confirmation comes on release.

Prepare storage: 594GB download + inference cache = plan for 1.2TB+ free storage minimum

vLLM readiness: Update vLLM to latest — KDA attention prefill cache support ships with the weights. Run pip install -U vllm before the drop.

If waiting for quants: Watch r/LocalLLaMA and Bartowski's HF profile — community GGUF Q4 builds typically appear within 12-24 hours of a major weight release

Check managed inference: Together AI, Fireworks AI, Groq — monitor their model pages from 8PM ET. Managed K3 inference on Western infrastructure eliminates the China NI Law concern without self-hosting overhead.

The Hallucination Finding — Read Before Deploying

Independent testing found a 51% hallucination rate that Moonshot AI omitted from its benchmark charts. This is a significant finding that is not disclosed in Moonshot's official K3 benchmark materials. The hallucination rate is particularly relevant for retrieval-augmented generation (RAG) pipelines, customer-facing applications, and any workload where factual accuracy is a hard requirement. K3's strengths — SWE Marathon #1 on agentic coding, Design Arena #1 on frontend coding — are real and independently verified. But the hallucination rate means K3 should be tested on factual accuracy benchmarks relevant to your use case before production deployment.

Hardware Reality — What Can Actually Run K3

SetupCan run K3?Notes
RTX 4090 / Mac StudioNoCannot load even Q4 — 1.4TB VRAM minimum for Q4
8× H100 80GBExperimentalMinimum load, very limited batch/context
18-24× H100 80GB (Q4)Production Q4~$50/hr reserved — practical deployment
64+ accelerators (full)Full qualityMoonshot recommended minimum for production

Why This Release Is Different From Other Open-Weight Drops

At 2.8 trillion parameters, Kimi K3 is the world's first 3T-class open-weights model. DeepSeek V4 Pro is approximately 1.6T parameters. Llama 3.1 405B is 405B. The scale difference is significant: K3 at full precision requires infrastructure that no previous open-weight model has required. The MXFP4 quantisation-aware training means Q4 quality degradation should be lower than models not trained this way — but the VRAM requirements remain extreme compared to any previous open-weight release. This is not a "run it on your gaming PC" release. It is a server-infrastructure release that happens to have open weights.

Sources: TechTimes · Wan27 · TechTimes (hallucination) · ExplainX · NodeMini · ByteIota · Related: Full self-hosting guide with cost estimates → · Kimi K3 full review → · Kimi K3 vs GPT-5.6 Sol →

Tags
KimiAI NewsGenerative AICoding AI2026

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →
💡 Kimi prompts
Prompt Guide
Best Kimi K3 Prompts for Coding (2026)
Kimi K3 ranks #1 globally on Design Arena frontend coding (1679 Elo) and #1 on SWE Marathon (42.0%) — making it the strongest available model for frontend development and long-horizon agentic coding tasks. At $3/$15/M with 1M context on Allegretto+ tier, it offers the best price-to-coding-capability ratio of any frontier model. These prompts are optimised for Kimi Code CLI, the Kimi API, and kimi.com on Allegretto+ or Vivace tiers.
Get Prompts →
Prompt Guide
Best Kimi K3 Prompts for Research (2026)
Kimi K3's 1M context window on Allegretto+ tier makes it one of the strongest research models available — it can ingest entire research papers, technical documentation sets, and large code repositories in a single session. At $0.30/M cached input, repeated context (like a standing research brief or knowledge base) becomes extremely cost-efficient. These prompts are optimised for kimi.com on Allegretto+ or Vivace tiers and the Kimi API.
Get Prompts →
Prompt Guide
Best Kimi K3 Prompts for Writing (2026)
Kimi K3 on Allegretto+ tier offers 1M context for writing tasks that require deep document awareness — editing a long manuscript, maintaining style consistency across a large content set, or rewriting with full context of everything written before. At $0.30/M cached input, repeatedly referencing a style guide or brand voice document costs almost nothing. These prompts are optimised for kimi.com on Allegretto+ and the Kimi API.
Get Prompts →