WED, JULY 22, 2026
Independent · In‑Depth · Practitioner‑Tested
Large Language Models

Gemini 3.5 Flash-Lite vs GPT-5.6 Luna (2026): $0.30/M vs $1/M — The Budget Tier Showdown

Google's Cheapest Option vs OpenAI's Budget Tier — High-Volume Routing at the Lowest Cost

🕐 6 min read 👁 16 views 📅 Jul 22, 2026

QUICK VERDICT — JULY 22, 2026

Cheapest option: Gemini 3.5 Flash-Lite — $0.30/$2.50/M vs Luna's $1/$6/M
Fastest throughput: Gemini 3.5 Flash-Lite — 350 tok/s vs Luna's ~150 tok/s
Better Terminal-Bench 2.1: GPT-5.6 Luna — 83.2% vs Flash-Lite's 54%
Better multi-step reasoning (Nerova): GPT-5.6 Luna — 41.3% vs Flash-Lite (not published on this benchmark)
Better long-context benchmark: Gemini 3.5 Flash-Lite — GDM-MRCR v2 72.2%
Use Flash-Lite for: agentic search fan-out, document processing, routing, classification — where speed and cost dominate
Use Luna for: tasks requiring stronger general reasoning at still-low cost

Full Comparison Table

ModelInput /1MOutput /1MSpeedTerminal-Bench 2.1Long-context
Gemini 3.5 Flash-Lite$0.30$2.50350 tok/s54%72.2% GDM-MRCR v2
GPT-5.6 Luna$1$6~150 tok/s83.2%Not published

Which to Choose

Gemini 3.5 Flash-Lite: Maximum cost efficiency for routing, classification, document extraction, and agentic search fan-out. At $0.30/$2.50/M it is 70-58% cheaper than Luna. 350 tok/s makes it the fastest option for high-concurrency workloads.

GPT-5.6 Luna: Better general reasoning (83.2% Terminal-Bench) for simple tasks that still require coherent multi-sentence outputs, Q&A, and light summarisation. The 29-point Terminal-Bench gap over Flash-Lite is significant for reasoning-dependent tasks despite higher cost.

Also consider: DeepSeek V4 Flash ($0.14/$0.28/M). If data residency is not a constraint, DeepSeek V4 Flash is still cheaper than both. Reminder: migrate deepseek-chat to deepseek-v4-flash before July 24 at 15:59 UTC.

Last updated July 22, 2026. Related: Full GPT-5.6 tier comparison → · Full Gemini July 21 launch → · DeepSeek V4 vs Sonnet 5 vs Terra →

⚖ Our Verdict

Gemini 3.5 Flash-Lite wins on price ($0.30/$2.50/M — 70% cheaper input, 58% cheaper output than Luna) and throughput (350 tok/s). GPT-5.6 Luna wins on Terminal-Bench 2.1 (83.2% vs Flash-Lite's 54%) and general reasoning quality. Use Flash-Lite for maximum throughput at minimum cost (routing, classification, agentic fan-out). Use Luna when reasoning quality at still-low cost matters.