💡 How to use these prompts:
Replace everything in [BRACKETS] with your specific details before sending.
Click Copy to copy any prompt to your clipboard instantly.
I want to measure what one unit of work actually costs me across different AI models.<br/><br/>Here is what my system does: [describe the workflow].<br/><br/>Help me define one "completed task" precisely enough to measure: where it starts, what has to be true for it to count as done, and what counts as a failure rather than a slow success.<br/><br/>Then list everything that consumes tokens between those two points - including retries, tool call round trips, system prompts re-sent on every turn, and context that gets resent as a conversation grows.<br/><br/>Tell me which of these I am most likely to forget to count.
Below are logs from my AI workflow, including failed and retried attempts.<br/><br/>Calculate: average tokens per successful task including all retries that preceded it, the share of total spend going to attempts that produced nothing, and how much a 10% improvement in first-attempt success would save.<br/><br/>Then tell me whether a more expensive model with a higher first-attempt rate would be cheaper overall, and what success rate it would need to break even.<br/><br/>LOGS:<br/>[paste]
My agent workflow makes tool calls. Here is a trace: [paste].<br/><br/>Work out how much of my token spend is the actual work versus the overhead of calling tools - the schemas, the results being read back into context, and the context resent on each subsequent turn.<br/><br/>Then tell me how much would change if the model batched calls instead of making them one at a time, and which of my tool calls could be batched without changing the result.
I am considering a premium speed tier that generates roughly [X] times faster at [Y] times the token cost.<br/><br/>Help me work out whether it pays. I need to account for: how long someone is actually blocked waiting, what that time is worth, whether the wait is on a critical path or in the background, and how many times a day this happens.<br/><br/>Give me the break-even - how many blocked minutes a day justify the multiplier - and tell me honestly which parts of my workload would see no benefit at all.<br/><br/>WORKLOAD:<br/>[describe]
I want to compare [model A] and [model B] on my own workload rather than on published benchmarks.<br/><br/>Design the test: how many runs for a meaningful result, which tasks to use and why those, what to hold constant, and what to measure besides cost - quality, first-attempt success, latency, failure modes.<br/><br/>Give me the table to fill in, with cost per completed task as the headline column.<br/><br/>Then tell me what result would make me switch, written now, before I see the numbers.<br/><br/>WORKLOAD:<br/>[describe]