I am pasting a sample of real requests from my application below, with their token counts if I have them.\n\nWork out the distribution, not the average. Specifically:\n\n1. What percentage of requests fall under 100,000 input tokens, and what percentage above.\n2. For the ones above, what is pushing them over. Separate system prompt, retrieved context, conversation history and user input as four distinct contributors.\n3. Which of those contributors is growing over the life of a session, and which is fixed.\n4. If I could move 20 percent of requests below the threshold, which change would do it with the least work.\n\nShow the maths. A five times price difference sits at that boundary, so I need to know how close I am, not roughly where I am.\n\n[PASTE SAMPLE REQUESTS]
Below are the distinct jobs my application sends to a language model, with example inputs and the outputs I need.\n\nPut each one in exactly one bucket and defend the choice:\n\nFRONTIER -- needs a top model. Name the specific capability and what fails without it.\nSMALL -- a cheap model handles this with a better prompt. Write that prompt.\nDETERMINISTIC -- does not need a model at all. Say what replaces it: a regex, a lookup, a rule.\n\nThen estimate monthly cost per bucket at [VOLUME] requests, using $0.10 in and $0.50 out per million for the small model and [YOUR CURRENT RATE] for the frontier one.\n\nDo not round in favour of using more model. If a job is borderline, say what test would settle it.\n\n[PASTE JOBS]
This prompt currently sends about [NUMBER] input tokens, which puts it over a pricing threshold at 100,000. I need it under, without losing output quality.\n\nAttack it in this order and show the token saving at each step:\n\n1. Remove instructions the model already follows reliably without being told.\n2. Replace pasted reference material with the smallest extract that still supports the answer.\n3. Cut examples down to the minimum number that holds the format. Say which ones are load-bearing.\n4. Move anything that is constant across requests into a cached prefix.\n5. Only then compress wording.\n\nGive me the rewritten prompt and an honest assessment of what quality risk each cut introduces. If getting under the threshold is not achievable without real loss, say so rather than shaving words.\n\n[PASTE PROMPT]
I want to route requests between a cheap model and an expensive one automatically. Here is what my application does.\n\nDesign the routing rule. Cover:\n\n- The signals available BEFORE the call that predict whether the cheap model will cope. Input length, task type, presence of certain fields.\n- The threshold for each signal, and how to pick it from data rather than intuition.\n- What happens on a miss: retry on the stronger model, or fail. Say which, and what that costs.\n- How to detect that the cheap model is quietly producing worse output rather than failing outright, which is the dangerous case.\n- What to log so the thresholds can be tuned later.\n\nGive me the decision logic as pseudocode, and tell me what percentage of traffic you expect to route each way.\n\n[PASTE APPLICATION DESCRIPTION]
Here is my system prompt and a description of how my application calls the model.\n\nCache reads cost a fraction of fresh input tokens, so anything stable that I am re-sending uncached is money burned. Tell me:\n\n1. Which parts of what I send are identical across requests and should sit in a cached prefix.\n2. Which parts look stable but are not, and would invalidate the cache. Be specific about what changes.\n3. How to reorder the prompt so the stable part comes first and the cache boundary falls in the right place.\n4. What my cache hit rate should be after this, and how to measure whether I got it.\n5. Where caching would NOT help, so I do not add complexity for nothing.\n\n[PASTE SYSTEM PROMPT AND CALL PATTERN]
I am considering moving [TASK] from [EXPENSIVE MODEL] to [CHEAPER MODEL]. Before I switch I need to know whether quality holds.\n\nBuild me an evaluation:\n\n- 20 test cases drawn from the examples below, chosen to cover the normal path, the edge cases and the ones where getting it wrong actually costs something.\n- For each, the specific property the output must have. Not good quality: a checkable property.\n- A scoring method I can run without a human reading every output.\n- The pass threshold, and your reasoning for setting it there.\n- The failure mode most likely to appear on a smaller model for this specific task, and the test case that would catch it.\n\nIf 20 cases is not enough to be confident for this task, say how many it would take.\n\n[PASTE EXAMPLES]