COST-PERFORMANCE RANKING — AUGUST 1, 2026
● #1 Price (input): Laguna S 2.1 — $0.10/M
● #2 Price: DeepSeek V4 Flash 0731 — $0.14/M
● #3 Price: Grok 4.5 — $2/M
● #4 Price: Kimi K3 API — $3/M (+ sanctions risk)
● #1 Terminal-Bench: Grok 4.5 — 83.3%
● #2 Terminal-Bench: V4 Flash 0731 — 82.7%
● #3 Terminal-Bench: Laguna S 2.1 — 70.2%
● #1 SWE Marathon: Kimi K3 — 42.0% (K3 not published on Terminal-Bench)
| Model | Input /1M | Terminal-Bench | Best for | Risk |
| Laguna S 2.1 | $0.10 | 70.2% | Lowest cost, DGX Spark deploy | None |
| DeepSeek V4 Flash 0731 | $0.14 | 82.7% | Best cost-performance, Codex-compatible | None |
| Grok 4.5 | $2.00 | 83.3% | Cursor, live X data | 54% hallucination (AA) |
| Kimi K3 | $3.00 | Not published | SWE Marathon #1 | Sanctions threat + 51% hallucination |
The new default recommendation for August 2026: DeepSeek V4 Flash 0731 at $0.14/M for API-first agentic coding workloads. It is 14× cheaper than Grok 4.5, 0.6 points behind on Terminal-Bench, open-weight MIT, and Codex-compatible out of the box. Laguna S 2.1 remains the choice where hardware accessibility (DGX Spark) or Western-company provenance are the deciding factors.
Last updated August 1, 2026. Related: V4 Flash 0731 full review → · Three coding agents comparison →