THU, OCTOBER 08, 2026
Independent · In‑Depth · Practitioner‑Tested
LLMs

Claude Haiku 5.5 vs GPT-6 Luna: Same Price, Different Math

Both list at $0.10 input and $0.50 output per million tokens. The difference is where the cheap band stops: 100,000 tokens on Haiku, 272,000 on Luna.

🕐 9 min read 👁 14 views 📅 Oct 8, 2026
THE SHORT ANSWER

● Under 100,000 prompt tokens — identical list price, $0.10 input and $0.50 output on both. Decide on quality, and Haiku 5.5 wins every benchmark row where both have a published number.

● Over 100,000 prompt tokens — Luna, clearly. Haiku jumps to $0.50 / $2.50. Luna holds $0.10 / $0.50 until 272,000.

● A 150,000-token prompt costs five times more on Haiku than on Luna. That one threshold is the entire decision.

The pricing, side by side

Per 1M tokensClaude Haiku 5.5GPT-6 Luna
Input, short prompts$0.10$0.10
Output, short prompts$0.50$0.50
Where the cheap band ends100,000 tokens272,000 tokens
Input, long prompts$0.50$0.20
Output, long prompts$2.50$0.75
Cached input$0.01Not published
Context window1,000,0001,050,000
Max output128,000128,000
Batch discount50%Not published

The 100K cliff, in money

Anthropic cut Haiku's headline price by 90% — Haiku 4.5 was $1.00 input and $5.00 output. The cut is real, and Anthropic says roughly 90% of Haiku 4.5 requests stayed under 100,000 tokens, so most users will only ever see the cheap rate.

The other 10% pay a lot more. Take a 150,000-token prompt returning 2,000 tokens:

  • Haiku 5.5: 150K at $0.50 = $0.075, plus 2K at $2.50 = $0.005. $0.080 per call.
  • GPT-6 Luna: 150K at $0.10 = $0.015, plus 2K at $0.50 = $0.001. $0.016 per call.

Five times the cost, for a prompt only 50% over the line. At a million calls a month that is $80,000 against $16,000. The pricing is not gradual — one token past 100,000 reprices the entire request, input and output both.

If your prompts cluster near 100K, measure before you migrate. A retrieval step that sometimes returns 95K tokens and sometimes 110K will produce a bill that looks broken.

Where Haiku earns the premium

On short prompts the price is a tie, so the benchmarks decide. Anthropic's launch table and public leaderboard snapshots put Haiku ahead on every row where both models report:

BenchmarkHaiku 5.5GPT-6 Luna
OSWorld 2.1 (offline, partial)72.4%48.9%
OSWorld 2.1 (strict pass)37.1%17.1%
Terminal-Bench 4.039.2%16.4%
FrontierCode 1.1 Main46.4%42.4%
Chartography (no tools)46.4%29.1%
GDPval-AA v2.1 (Elo)16201437
AA-Briefcase v1.1 (Elo)15781336

The computer-use gap is the one that matters in practice. 72.4% against 48.9% on OSWorld, and 39.2% against 16.4% on Terminal-Bench, is not a rounding difference — it is the difference between an agent step that usually works and one that usually does not. For context, Haiku 4.5 scored 0.0% on Terminal-Bench 4.0.

Two caveats worth stating. Most of these are Anthropic's own published figures, and a vendor's launch table is not an independent test. And Haiku's Humanity's Last Exam score of 45.9% without tools, 57.4% with, is a reminder that this is a small model: Sonnet 5.5 leads every row, including 70.6% on Terminal-Bench.

Which one to pick

  • Short prompts, agent steps, computer use → Haiku 5.5. Same price, materially better scores on the tasks where a small model is actually deployed.
  • Long documents, large retrieval contexts, anything past 100K → GPT-6 Luna. The pricing is not close.
  • Prompt length varies unpredictably → Luna. Predictable billing beats a slightly better model you cannot forecast the cost of.
  • Heavy repeated system prompts → Haiku 5.5. Cached input at $0.01 per million is published; Luna's is not.
  • Coding and terminal work above both → GPT-6 Sol at $2.00 / $10.00, which scores 49.3% on FrontierCode and 49.4% on Terminal-Bench. Twenty times the price, so use it as an escalation path rather than a default.

The stack that actually works

Nobody serious runs one model. The arrangement that uses these prices properly:

  1. Haiku 5.5 for classification, extraction and short agent steps — the high-volume layer, where the $0.10 rate applies and the benchmark lead is real.
  2. GPT-6 Luna for anything that ingests a long document, so the 100K cliff never fires.
  3. GPT-6 Sol or Claude Sonnet 5.5 as the escalation tier when the cheap layer returns low confidence.
  4. Prompt caching on the Haiku layer at $0.01 per million cached input. A 20,000-token system prompt reused across a million calls is the single largest line item most teams never look at.

FAQ

Is Claude Haiku 5.5 really 90% cheaper?
On short prompts, yes. Haiku 4.5 cost $1.00 input and $5.00 output per million; Haiku 5.5 is $0.10 and $0.50. Above 100,000 prompt tokens it is $0.50 and $2.50, which is a 50% cut rather than 90%.
Does the 100K threshold count output tokens too?
The threshold is measured on the prompt. Once a request crosses it, both the input and the output for that request are billed at the higher rate.
Which is better for AI agents?
Haiku 5.5, on the published evidence. It scores 72.4% on OSWorld 2.1 against Luna's 48.9%, and 39.2% on Terminal-Bench 4.0 against 16.4%. Agent steps are usually short, so the pricing tie holds.
What is the context window on each?
Haiku 5.5 is 1,000,000 tokens with 128,000 max output, rising to 300,000 output for batch jobs in beta. GPT-6 Luna is 1,050,000 with 128,000 max output. Both can accept far more than their cheap pricing band covers.
Is there anything cheaper than either?
Gemini 2.5 Flash-Lite lists at $0.10 input and $0.40 output with no threshold, so it is marginally cheaper on output at any length — but it is an older generation and does not compete on the agentic benchmarks above. DeepSeek V4.1 Flash is $0.30 / $1.20 at peak and half that off-peak.

Sources

Full launch detail and the complete benchmark table: Claude Haiku 5.5 pricing and benchmarks.

⚖ Our Verdict

Tie on list price under 100K tokens, and Haiku 5.5 wins every published benchmark — so Haiku is the better buy for short prompts and agent steps. Past 100,000 tokens Luna costs a fifth as much, and no benchmark gap closes a 5x price difference.