Here are my agent logs and the outcome of each run.<br/><br/>Do not give me cost per million tokens. Give me cost per pull request that was actually MERGED.<br/><br/>Break out: total spend, runs that produced a mergeable PR, runs that produced a PR that was rejected or rewritten, and runs that produced nothing.<br/><br/>Then tell me what share of total spend went to the last two categories. That is the number I want.<br/><br/>LOGS:<br/>[paste]
Below are my agent runs with the size of the change produced and the tokens consumed.<br/><br/>Test whether token spend correlates with output size or with something else - number of files touched, test failures encountered, retries, time spent.<br/><br/>Tell me which variable actually predicts cost, and give me a formula I can use to estimate a task's cost before running it.<br/><br/>If the data does not support a clean predictor, say so rather than fitting one.<br/><br/>RUNS:<br/>[paste]
From the task history below, identify categories where the agent's success rate is low enough that supervised work would have been cheaper.<br/><br/>For each category give: attempt count, success rate, average cost per attempt, average cost per success, and an estimate of what a developer would have cost.<br/><br/>Recommend which categories to stop sending to the agent entirely. Be specific - I would rather cut three categories than shave 5% off everything.<br/><br/>HISTORY:<br/>[paste]
My current agent inference cost is below. OpenAI has disclosed that continuous behavioural monitoring costs about 20% additional inference compute.<br/><br/>Recalculate my cost per completed task with that overhead applied, and show the break-even against a human doing the same work.<br/><br/>Then tell me at what success rate the agent stops being cheaper than the human, with monitoring included.<br/><br/>COSTS:<br/>[paste]
My agent runs on the model below at current list prices.<br/><br/>Model three scenarios: prices rise 2x, 3x, and 4.5x - the range DeepSeek actually moved this year.<br/><br/>For each: my new monthly cost, whether the deployment is still cheaper than the human alternative, and what I would have to change to stay viable.<br/><br/>Then tell me which of those changes I should make now regardless, because they are good ideas at current prices too.<br/><br/>USAGE:<br/>[paste]
For each completed agent task below, I have recorded how long a human spent reviewing the output.<br/><br/>Calculate the fully-loaded cost per task including review time at my engineering hourly rate, and compare it against the cost of that engineer doing the task themselves.<br/><br/>Flag any category where review time is so high that the agent is producing negative value.<br/><br/>Asynchronous agents shift work from doing to reviewing - I want to know where that trade stopped being worth it.<br/><br/>TASKS AND REVIEW TIMES:<br/>[paste]
Using everything above, write a one-page summary of my agent deployment's economics for someone who does not work in engineering.<br/><br/>Include: total spend, tasks completed, cost per completed task, the human-equivalent cost, net position, and the three largest sources of waste.<br/><br/>Rules: no claim without a number behind it, state uncertainty where the data is thin, and do not round in my favour.<br/><br/>If the deployment is not paying for itself, lead with that.<br/><br/>DATA:<br/>[paste]