QUICK VERDICT — AUGUST 2026
● Price (input): V4 Flash — $0.14/M vs K3 $3/M (21× cheaper)
● Terminal-Bench 2.1: V4 Flash — 82.7%. K3 — not published on this benchmark.
● SWE Marathon: Kimi K3 — 42.0% #1 globally. V4 Flash — not published.
● Intelligence index: K3 — BenchLM #5 of 214, AA Index top-5. V4 Flash — agent specialist, not general frontier.
● Open weights: Both — K3 Modified MIT (2.8T, 8× H100 min). V4 Flash MIT (284B, smaller hardware).
● Hallucination: K3 — 51% in independent testing. V4 Flash — not independently tested yet.
● Regulatory risk: K3 — US Treasury sanctions warning on Moonshot. V4 Flash — none active.
● Codex-compatible: V4 Flash — native. K3 — OpenAI-compatible API but not Codex-specific.
What Each Model Is Actually Built For
DeepSeek V4 Flash 0731 and Kimi K3 are frontier-competitive models that serve different primary use cases. According to DeepSeek's official changelog, V4 Flash 0731 was specifically re-post-trained on agent-focused data — its 82.7% Terminal-Bench score, 54.4% DeepSWE score, and 76.7% Cybergym score are all agent-benchmark results. It is a coding and terminal execution specialist. Kimi K3 is a general intelligence frontier model that happens to lead on SWE Marathon: its BenchLM #5 ranking covers general reasoning, multimodal, and coding, not just agent tasks. Choosing between them means first deciding whether your workload is agent/terminal-execution-focused or general-intelligence-focused.
The size difference reinforces this. V4 Flash 0731 is 284B total, 13B active per token — it runs on accessible hardware and is the cheapest capable agentic model per independent benchmark. Kimi K3 is 2.8T total, 104B active — it requires a minimum of 8× H100 80GB to load and produces higher-quality outputs on reasoning tasks requiring deep context and multi-step judgment. As Kingy.ai's economics analysis notes, the 21× input price gap between the two models represents a material operational decision at volume — not a minor cost line.
Benchmark Comparison
| Benchmark | DeepSeek V4 Flash 0731 | Kimi K3 |
| Terminal-Bench 2.1 | 82.7% | Not published |
| SWE Marathon | Not published | 42.0% #1 |
| DeepSWE | 54.4% | Not published |
| BenchLM Intelligence Index | Not ranked (specialist) | #5 of 214 |
| Price (input / output) | $0.14 / $0.28/M | $3 / $15/M |
Decision Framework
DeepSeek V4 Flash 0731 for:
High-volume agentic coding loops where Terminal-Bench 82.7% is sufficient and $0.14/M vs $3/M changes the economics of your deployment. Codex-compatible out of the box, MIT weights, no regulatory risk. The cheapest capable agentic model with published benchmarks as of August 2026. Migrate from deepseek-chat before October 24, 2026.
Kimi K3 for:
Long-horizon multi-file agentic coding (SWE Marathon #1), general reasoning at frontier level (BenchLM #5), and workloads where the intelligence ceiling matters more than the 21× cost premium. Self-hosted K3 on Western cloud via Modified MIT weights reduces the sanctions risk and the China data residency concern. Test hallucination rate (51% in independent testing) on your specific task type before production deployment.
Kimi K3 regulatory note:
US Treasury Secretary Bessent's July 2026 warning on potential Moonshot Entity List designation is unresolved as of August 2026. Freeze new Moonshot-hosted API commitments until resolved. Self-hosted K3 via Modified MIT weights on Western cloud infrastructure is the lower-risk enterprise path.
Sources: DeepSeek official changelog · HuggingFace model card · Kingy.ai K3 economics · Related: DeepSeek V4 Flash 0731 full review → · Kimi K3 open weights guide → · DeepSeek V4 Flash vs Laguna S 2.1 →