THU, OCTOBER 08, 2026
Independent · In‑Depth · Practitioner‑Tested
Claude Coding

AI Cost Optimization Prompts: 6 to Cut an API Bill

Claude Haiku 5.5 arrived on 7 October 2026 at $0.10 per million input tokens below 100,000 tokens, and $0.50 above it. GPT-6 Luna is priced identically. That five-times cliff at a token threshold is the single biggest lever on most API bills right now, and almost nobody knows where their own prompts sit relative to it. These six prompts work through the audit in order: measure what you are actually sending, find what does not need a frontier model, rewrite what sits just over a tier boundary, then build the eval that tells you whether the cheaper model held up. Run them against your real logs, not a sample you tidied.

⌨️ 6 prompts 🕐 Updated Oct 8, 2026
💡 How to use these prompts: Replace everything in [BRACKETS] with your specific details before sending. Click Copy to copy any prompt to your clipboard instantly.
1
Find out where your prompts actually sit
Knowing which side of a pricing tier your traffic falls on
I am pasting a sample of real requests from my application below, with their token counts if I have them.\n\nWork out the distribution, not the average. Specifically:\n\n1. What percentage of requests fall under 100,000 input tokens, and what percentage above.\n2. For the ones above, what is pushing them over. Separate system prompt, retrieved context, conversation history and user input as four distinct contributors.\n3. Which of those contributors is growing over the life of a session, and which is fixed.\n4. If I could move 20 percent of requests below the threshold, which change would do it with the least work.\n\nShow the maths. A five times price difference sits at that boundary, so I need to know how close I am, not roughly where I am.\n\n[PASTE SAMPLE REQUESTS]
2
Sort the workload by what it genuinely needs
Deciding what to move off an expensive model
Below are the distinct jobs my application sends to a language model, with example inputs and the outputs I need.\n\nPut each one in exactly one bucket and defend the choice:\n\nFRONTIER -- needs a top model. Name the specific capability and what fails without it.\nSMALL -- a cheap model handles this with a better prompt. Write that prompt.\nDETERMINISTIC -- does not need a model at all. Say what replaces it: a regex, a lookup, a rule.\n\nThen estimate monthly cost per bucket at [VOLUME] requests, using $0.10 in and $0.50 out per million for the small model and [YOUR CURRENT RATE] for the frontier one.\n\nDo not round in favour of using more model. If a job is borderline, say what test would settle it.\n\n[PASTE JOBS]
3
Rewrite a prompt to fit under the cliff
Getting a request under a token threshold without degrading it
This prompt currently sends about [NUMBER] input tokens, which puts it over a pricing threshold at 100,000. I need it under, without losing output quality.\n\nAttack it in this order and show the token saving at each step:\n\n1. Remove instructions the model already follows reliably without being told.\n2. Replace pasted reference material with the smallest extract that still supports the answer.\n3. Cut examples down to the minimum number that holds the format. Say which ones are load-bearing.\n4. Move anything that is constant across requests into a cached prefix.\n5. Only then compress wording.\n\nGive me the rewritten prompt and an honest assessment of what quality risk each cut introduces. If getting under the threshold is not achievable without real loss, say so rather than shaving words.\n\n[PASTE PROMPT]
4
Design the routing policy
Splitting traffic between models without hand-picking each call
I want to route requests between a cheap model and an expensive one automatically. Here is what my application does.\n\nDesign the routing rule. Cover:\n\n- The signals available BEFORE the call that predict whether the cheap model will cope. Input length, task type, presence of certain fields.\n- The threshold for each signal, and how to pick it from data rather than intuition.\n- What happens on a miss: retry on the stronger model, or fail. Say which, and what that costs.\n- How to detect that the cheap model is quietly producing worse output rather than failing outright, which is the dangerous case.\n- What to log so the thresholds can be tuned later.\n\nGive me the decision logic as pseudocode, and tell me what percentage of traffic you expect to route each way.\n\n[PASTE APPLICATION DESCRIPTION]
5
Audit the caching
Making sure the cheap cached rate applies to everything it could
Here is my system prompt and a description of how my application calls the model.\n\nCache reads cost a fraction of fresh input tokens, so anything stable that I am re-sending uncached is money burned. Tell me:\n\n1. Which parts of what I send are identical across requests and should sit in a cached prefix.\n2. Which parts look stable but are not, and would invalidate the cache. Be specific about what changes.\n3. How to reorder the prompt so the stable part comes first and the cache boundary falls in the right place.\n4. What my cache hit rate should be after this, and how to measure whether I got it.\n5. Where caching would NOT help, so I do not add complexity for nothing.\n\n[PASTE SYSTEM PROMPT AND CALL PATTERN]
6
Build the eval before you switch
Proving the cheaper model is good enough before committing
I am considering moving [TASK] from [EXPENSIVE MODEL] to [CHEAPER MODEL]. Before I switch I need to know whether quality holds.\n\nBuild me an evaluation:\n\n- 20 test cases drawn from the examples below, chosen to cover the normal path, the edge cases and the ones where getting it wrong actually costs something.\n- For each, the specific property the output must have. Not good quality: a checkable property.\n- A scoring method I can run without a human reading every output.\n- The pass threshold, and your reasoning for setting it there.\n- The failure mode most likely to appear on a smaller model for this specific task, and the test case that would catch it.\n\nIf 20 cases is not enough to be confident for this task, say how many it would take.\n\n[PASTE EXAMPLES]