QUICK VERDICT — JULY 22, 2026
● Cheapest option: Gemini 3.5 Flash-Lite — $0.30/$2.50/M vs Luna's $1/$6/M
● Fastest throughput: Gemini 3.5 Flash-Lite — 350 tok/s vs Luna's ~150 tok/s
● Better Terminal-Bench 2.1: GPT-5.6 Luna — 83.2% vs Flash-Lite's 54%
● Better multi-step reasoning (Nerova): GPT-5.6 Luna — 41.3% vs Flash-Lite (not published on this benchmark)
● Better long-context benchmark: Gemini 3.5 Flash-Lite — GDM-MRCR v2 72.2%
● Use Flash-Lite for: agentic search fan-out, document processing, routing, classification — where speed and cost dominate
● Use Luna for: tasks requiring stronger general reasoning at still-low cost
Full Comparison Table
| Model | Input /1M | Output /1M | Speed | Terminal-Bench 2.1 | Long-context |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 350 tok/s | 54% | 72.2% GDM-MRCR v2 |
| GPT-5.6 Luna | $1 | $6 | ~150 tok/s | 83.2% | Not published |
Which to Choose
Gemini 3.5 Flash-Lite: Maximum cost efficiency for routing, classification, document extraction, and agentic search fan-out. At $0.30/$2.50/M it is 70-58% cheaper than Luna. 350 tok/s makes it the fastest option for high-concurrency workloads.
GPT-5.6 Luna: Better general reasoning (83.2% Terminal-Bench) for simple tasks that still require coherent multi-sentence outputs, Q&A, and light summarisation. The 29-point Terminal-Bench gap over Flash-Lite is significant for reasoning-dependent tasks despite higher cost.
Also consider: DeepSeek V4 Flash ($0.14/$0.28/M). If data residency is not a constraint, DeepSeek V4 Flash is still cheaper than both. Reminder: migrate deepseek-chat to deepseek-v4-flash before July 24 at 15:59 UTC.
Last updated July 22, 2026. Related: Full GPT-5.6 tier comparison → · Full Gemini July 21 launch → · DeepSeek V4 vs Sonnet 5 vs Terra →