SAT, SEPTEMBER 26, 2026
Independent · In‑Depth · Practitioner‑Tested
Claude Productivity

Is Your AI Optimising the Metric or the Outcome? Seven Audit Prompts

Blue Cross found AI scribes added $942 million in costs. Nobody asked the scribes to increase billing - they were asked to document thoroughly, and thorough documentation is what triggers higher payment. The tell was that anemia diagnoses rose while blood transfusions did not. These seven prompts run that same check on your own deployment: is the AI producing more of the measured thing, or more of the valuable thing?

⌨️ 7 prompts 🕐 Updated Sep 24, 2026
💡 How to use these prompts: Replace everything in [BRACKETS] with your specific details before sending. Click Copy to copy any prompt to your clipboard instantly.
1
Separate the Metric From the Outcome
The whole Blue Cross finding in one question. Documenting thoroughly and billing more were the same action.
Here is what my AI deployment does and how we measure it.<br/><br/>Answer two questions separately and explicitly:<br/>1. What does this system actually optimise? Not what we want it to do - what does the measurement reward?<br/>2. What is the outcome we actually care about?<br/><br/>Then state plainly where those two come apart, and what behaviour would score well on 1 while being useless or harmful for 2.<br/><br/>Do not reassure me that they are aligned unless you can show it.<br/><br/>DEPLOYMENT:<br/>[paste]
2
The Diagnosis-Without-Treatment Test
Anemia diagnoses rose; transfusions did not. That gap is the entire case and it generalises.
My AI produces more of X since deployment. I want to know whether that is real or artefact.<br/><br/>Tell me what downstream thing should ALSO have increased if the rise in X were genuine.<br/><br/>Be specific and name something independently measurable that nobody would bother to game.<br/><br/>Then tell me how to check it, and what it means if X rose and the downstream thing did not.<br/><br/>WHAT X IS:<br/>[paste]
3
Where Is the Money Attached
Incentive drift needs no bad actor. Find the points where output and payment touch.
Map every point in this workflow where a number determines a payment, a budget, a bonus, a ranking, or a headcount.<br/><br/>For each one, state whether the AI can influence that number, and by how much.<br/><br/>Sort by how much money moves per unit of influence.<br/><br/>I am not asking whether anyone intends to game it. I am asking where it would happen without anyone intending anything.<br/><br/>WORKFLOW:<br/>[paste]
4
Reconstruct the Baseline
The $942m figure exists only because there was a 2023 comparison point. Build yours retrospectively if you must.
I deployed this AI system without recording a proper before-state.<br/><br/>From the data below, reconstruct the best baseline you can for the period before deployment.<br/><br/>Be explicit about what you can and cannot establish. Where a comparison is not valid because something else changed at the same time, say so rather than producing a number.<br/><br/>A defensible "we cannot tell" is more useful to me than a confident figure I will quote in a meeting.<br/><br/>DATA:<br/>[paste]
5
Steelman the Benign Explanation
Some of the Blue Cross increase is genuine under-coding being fixed. Find out which part yours is.
Here is a pattern that looks like my AI is gaming a metric.<br/><br/>Argue the other side as strongly as you can. What is the legitimate explanation in which this rise is real, correct and good?<br/><br/>Then tell me what evidence would distinguish the benign explanation from the gaming one.<br/><br/>Do not hedge between them. Give me the discriminating test.<br/><br/>PATTERN:<br/>[paste]
6
Audit the Vendor's Claim
"Documents more thoroughly" and "increases reimbursement" were the same feature sold twice.
Below is a vendor's claim about what their AI tool delivers.<br/><br/>For each claimed benefit, tell me:<br/>1. What exactly is being measured<br/>2. Who benefits if that number goes up<br/>3. Whether the claim would still hold if the number rose without the underlying reality changing<br/><br/>Flag every claim where the metric and the benefit are the same measurement wearing two names.<br/><br/>CLAIM:<br/>[paste]
7
Write the Monitoring That Would Catch It
Blue Cross caught this over two years. A monthly check catches it in one.
Design the smallest ongoing check that would tell me if this AI system starts optimising the metric instead of the outcome.<br/><br/>Requirements:<br/>1. It must measure a downstream reality, not the system's own output<br/>2. It must be cheap enough to run monthly forever<br/>3. It must have a threshold defined NOW, before I have any emotional stake in the result<br/>4. State what I will do if the threshold trips<br/><br/>One page. If I cannot run it every month, it is too complicated.<br/><br/>SYSTEM:<br/>[paste]