THREE-WAY CODING AGENT COMPARISON — JULY 30, 2026
● Terminal-Bench 2.1: Grok Build 83.3% > Laguna 70.2% > Kimi (not published)
● SWE Marathon: Kimi Code 42.0% #1 > Grok Build / Laguna (not published)
● API cost: Laguna $0.10/M > Grok Build $2/M > Kimi Code $3/M
● Open weights: Laguna (OpenMDW) · Kimi (Modified MIT) · Grok Build (closed model, Apache 2.0 harness)
● Regulatory risk: Kimi Code (Treasury sanctions threat) · others none
● Privacy incident: Grok Build (SSH key upload, patched July 25, open-sourced)
Grok Build for: Terminal task execution where Terminal-Bench 83.3% matters and $2/M is acceptable. Available on Cursor (all plans). X/live data access if relevant to your workflow. If you used Grok Build before July 25 patch: rotate any credentials from directories you ran it in.
Laguna S 2.1 for: Cost-sensitive high-volume agentic coding where $0.10/M vs $2-3/M matters at scale. Self-hostable on DGX Spark. No regulatory risk. Terminal-Bench 70.2% is competitive even against much larger models. The new default recommendation for teams starting fresh with open-weight coding in July 2026.
Kimi Code for: Long-horizon agentic tasks where SWE Marathon #1 performance is specifically what you need. Freeze new Moonshot API commitments until Treasury sanctions situation resolves. Self-hosted K3 on Western cloud (via Modified MIT weights) is lower risk than Moonshot API.
Last updated July 30, 2026. Related: Laguna S 2.1 full review → · Grok 4.5 / Grok Build review → · Kimi Code vs Codex →