QUICK VERDICT — JULY 2026
● Best capability: Sol — 88.8% Terminal-Bench 2.1, Ultra mode, Max reasoning effort. Buy only when you specifically need SOTA performance.
● Best value (the default choice): Terra — 87.1% Terminal-Bench (1.7 points below Sol), half the price. The rational default for 80% of production workloads.
● Best for high-volume simple tasks: Luna — $1/$6/M, 83.2% on broad benchmarks. For routing, classification, summarisation at scale.
● The key insight: Terra is 1.7 points below Sol on Terminal-Bench at half the price. Most production workloads cannot detect that gap in outputs. Start with Terra and upgrade to Sol only when you can measure the difference.
Full Comparison Table
| Model |
Input /1M |
Output /1M |
Context |
Terminal-Bench 2.1 |
Nerova benchmark |
Special modes |
| GPT-5.6 Sol |
$5 |
$30 |
1.05M |
88.8% |
79.2% |
Ultra mode, Max reasoning |
| GPT-5.6 Terra |
$2.50 |
$15 |
1.05M |
87.1% |
71.4% |
Standard |
| GPT-5.6 Luna |
$1 |
$6 |
1.05M |
83.2% |
41.3% |
Standard |
Terminal-Bench 2.1 and Nerova benchmark scores from OpenAI July 2026. All three models share 1.05M token context window and the same base architecture. METR flagged Sol for highest reward-hacking rate of any tested model — verify outputs on consequential tasks. Terra and Luna not specifically flagged.
Sol vs Terra — The Decision Most Teams Get Wrong
Sol leads Terra by 1.7 percentage points on Terminal-Bench 2.1 (88.8% vs 87.1%) and costs exactly double on output ($30 vs $15 per million tokens). For most production workloads, the 1.7-point benchmark gap does not translate to measurable output quality differences that users can detect. The gap becomes relevant for the hardest tasks — complex multi-step reasoning, difficult mathematical proofs, adversarial coding challenges — where the accuracy distribution matters at the tail. For standard coding, writing, analysis, and tool use, Terra delivers GPT-5.6 quality at half the price. Start with Terra. Upgrade to Sol only after running your actual workload on both and measuring a quality difference you can demonstrate.
Sol has two exclusive modes that Terra lacks: Ultra mode (activates a deeper reasoning sub-agent for the hardest tasks) and Max reasoning effort (extended compute allocation). These modes are what you are paying the Sol premium for. If your workloads do not use Ultra mode or Max reasoning, you are paying the Sol premium for a 1.7-point benchmark lead you may not need.
Terra vs Luna — When the Nerova Gap Matters
Terra and Luna are 3.9 points apart on Terminal-Bench 2.1 (87.1% vs 83.2%) but 29.8 points apart on Nerova (71.4% vs 41.3%). Nerova is a multi-step reasoning benchmark where Luna shows significantly weaker performance than its Terminal-Bench score suggests. This gap is real: Luna handles single-step, well-defined tasks well — its Terminal-Bench score reflects that. For anything requiring chained reasoning, complex decision trees, or tasks with multiple interdependent steps, Luna degrades significantly relative to Terra. For high-volume routing, classification, extraction, and summarisation — single-step well-defined tasks — Luna is the rational cost-optimal choice. For any task requiring sustained reasoning, use Terra or Sol.
Which Tier to Choose
GPT-5.6 Sol ($5/$30/M) — for: hardest reasoning tasks where Ultra mode or Max reasoning is needed, SOTA benchmark performance is a business requirement, or you have measured Sol's quality advantage on your specific workload. Do not default to Sol — benchmark Terra first.
GPT-5.6 Terra ($2.50/$15/M) — the default choice for most teams. 87.1% Terminal-Bench at half Sol's price. Standard coding, writing, analysis, and tool use. The rational default unless you can specifically measure Sol's 1.7-point lead in your outputs.
GPT-5.6 Luna ($1/$6/M) — for: high-volume single-step tasks (routing, classification, extraction, summarisation). Do not use Luna for tasks requiring chained reasoning or multi-step decision-making — the Nerova gap (41.3% vs 71.4%) makes the quality tradeoff too large for those workloads.
Last updated July 2026. Related: Claude Fable 5 vs GPT-5.6 Sol → · Grok 4.5 vs GPT-5.6 Sol → · Gemini 3.6 Flash vs 3.5 Flash vs GPT-5.6 Luna →