THE SHORT ANSWER
● Under 100,000 tokens → Haiku 5.5. $0.10 / $0.50 against $0.30 / $2.50. Three times cheaper on input, five times on output. Not close.
● Over 100,000 tokens → Gemini 3.5 Flash-Lite. Haiku reprices to $0.50 / $2.50; Flash-Lite never changes rate.
● Generating more than 65,536 tokens → Haiku 5.5, whatever the prompt length. Flash-Lite caps output at half of Haiku's 128,000.
The pricing, side by side
| Per 1M tokens | Claude Haiku 5.5 | Gemini 3.5 Flash-Lite |
| Input, under 100K | $0.10 | $0.30 |
| Output, under 100K | $0.50 | $2.50 |
| Input, over 100K | $0.50 | $0.30 |
| Output, over 100K | $2.50 | $2.50 |
| Cached input | $0.01 | $0.03 |
| Batch rate | $0.05 / $0.25 | $0.15 / $1.25 |
| Context window | 1,000,000 | 1,048,576 |
| Max output | 128,000 | 65,536 |
What the numbers mean in practice
For the common case — a few thousand tokens in, a few hundred out — Haiku is dramatically cheaper. A million calls at 2,000 input and 300 output tokens:
- Haiku 5.5: $200 input, $150 output. $350.
- Gemini 3.5 Flash-Lite: $600 input, $750 output. $1,350.
Nearly four times the bill for the same work. At batch rates the gap widens: $0.05 / $0.25 against $0.15 / $1.25.
Push the same comparison to a 200,000-token prompt and it inverts. Haiku bills that request at $0.50 input and $2.50 output; Flash-Lite stays at $0.30 and $2.50. Flash-Lite becomes the cheaper option the moment you cross 100,000 tokens, and stays cheaper all the way to the context limit.
The crossover sits exactly on Haiku's threshold. There is no window where the two are close — one is clearly cheaper on either side of 100,000 tokens.
The output limit nobody checks
Flash-Lite maxes out at 65,536 output tokens. Haiku 5.5 does 128,000, and up to 300,000 for batch jobs in beta.
For classification or extraction this never comes up. For generating a long document, translating a book chapter, or producing a large structured JSON payload in one call, it is a hard wall rather than a price difference — and it is the kind of limit that surfaces in production rather than in testing, because test inputs are usually short.
Which one to pick
- High-volume short tasks → Haiku 5.5. Classification, extraction, routing, summarising. Roughly a quarter of the cost.
- Long-document ingestion → Gemini 3.5 Flash-Lite. Past 100K tokens the flat rate wins, and it keeps winning.
- Long generation → Haiku 5.5. Twice the output ceiling.
- Already inside Google Cloud → Flash-Lite, unless volume is high enough that a 4x difference outweighs the integration work.
- Tightest possible budget → look at Gemini 2.5 Flash-Lite, $0.10 input and $0.40 output flat, with no threshold. Older generation and weaker on agentic work, but genuinely the cheapest flat rate on this list.
Where each fits in a stack
- Haiku 5.5 as the default high-volume worker on short prompts.
- Flash-Lite on the one path that ingests whole documents, so Haiku's cliff never fires.
- Prompt caching on whichever handles the repeated system prompt — $0.01 per million on Haiku, $0.03 on Flash-Lite.
- Route by prompt length, not by model preference. A single token count at the edge of your pipeline picks the cheaper model per request, and it is roughly ten lines of code.
FAQ
Which is cheaper overall?
Haiku 5.5, for almost everyone. It is three times cheaper on input and five times on output below 100,000 prompt tokens, and Anthropic says roughly 90% of requests to the previous Haiku stayed under that line.
Does Gemini 3.5 Flash-Lite have a long-context surcharge?
No. Google applies a higher meter to Pro-tier models above 200,000 tokens, but the Flash and Flash-Lite rates are flat at any prompt length.
What about Gemini 3.5 Flash, not Lite?
$1.50 input and $9.00 output per million — fifteen and eighteen times Haiku 5.5's short-prompt rate. It is a different weight class and not a competitor to Haiku on cost.
Which has the bigger context window?
Effectively a tie: 1,000,000 against 1,048,576. The meaningful difference is output, where Haiku allows 128,000 tokens and Flash-Lite 65,536.
Sources
Full launch detail: Claude Haiku 5.5 pricing and benchmarks.