MON, OCTOBER 05, 2026
Independent · In‑Depth · Practitioner‑Tested
Claude Productivity

Six Prompts to Read a Benchmark Before You Repeat It

MoralityBench published a leaderboard this week ranking 13 models on adapted moral psychology questionnaires. Its own documentation states that the metrics "do not measure moral correctness or deployment safety", and that human validation of the original instruments does not establish validity of the AI adaptation. It is going to be cited as a morality ranking regardless. The gap between what an eval measures and what it gets quoted as is where most confident wrong claims about AI come from, and it is almost always documented. These six prompts find it.

⌨️ 6 prompts 🕐 Updated Oct 5, 2026
💡 How to use these prompts: Replace everything in [BRACKETS] with your specific details before sending. Click Copy to copy any prompt to your clipboard instantly.
1
Find What the Authors Say It Does Not Measure
The caveat is almost always published and almost never quoted. Reading it first changes what you are willing to claim.
Here is a benchmark, eval or study I am about to cite: [paste the paper, repo README, methodology page or leaderboard notes].<br/><br/>Find and quote exactly:<br/>- what the authors state it measures<br/>- what they state it does NOT measure<br/>- every limitation or caveat they put in writing themselves<br/><br/>Put the authors' own disclaimers first, before any summary of results. If they describe it as exploratory, preliminary or unvalidated, lead with that.<br/><br/>If there is no methodology documentation at all, say so - that is the finding.
2
Work Out What the Reference Point Actually Is
"Closest to human norms" usually means closest to a survey mean. That is a real finding and not the one the headline implies.
For this benchmark: [paste], tell me precisely what the scores are measured AGAINST.<br/><br/>Is it a human baseline, and if so which humans, sampled how, when? Is it another model? An absolute standard? A survey average?<br/><br/>Then tell me what "scoring well" literally means in those terms - not what it sounds like it means.<br/><br/>If the reference is a population average, say plainly that closer to average is what is being rewarded, and what that does and does not imply.
3
Check Whether the Instrument Was Modified
Four changed items can be the difference between a validated measure and an unvalidated one wearing its name.
This evaluation adapts an existing instrument, test or dataset: [describe or paste].<br/><br/>Tell me what was changed - items replaced, wording altered, scales rescored, prompts restructured - and what the original validation covered.<br/><br/>Then tell me honestly whether the original validation still applies after those changes, and which specific claims it no longer supports.<br/><br/>Do not be generous about this. A validated instrument with altered items is a new instrument.
4
Find the Reliability Number Before the Ranking
A score with no reliability figure is a single measurement dressed as a finding. Consistency usually matters more than placement.
Here are a benchmark's results: [paste].<br/><br/>Before interpreting any rank, tell me:<br/>- how many runs per model<br/>- what the variance or self-consistency across runs was<br/>- whether error bars, confidence intervals or agreement rates are reported at all<br/><br/>Then tell me which differences between adjacent ranks are smaller than the run-to-run variation, and therefore are not differences.<br/><br/>If no reliability figure is reported, say that the ranking is one sample and should be read as such.
5
Rewrite the Headline So It Is True
Useful before you publish, quote or forward anything. Half the time the honest version is still interesting.
Here is how this result is being reported: [paste the headline, post or summary].<br/>Here is what the benchmark actually measures: [paste from the earlier prompts].<br/><br/>Rewrite the claim so it is accurate and still worth saying. Keep it short.<br/><br/>Then tell me which parts of the original are overstatement, which are category errors, and which are simply wrong.<br/><br/>If the accurate version is boring, say so - that tells me whether the result was worth citing at all.
6
Decide Whether It Should Change What You Do
The final filter. A benchmark can be sound, widely cited, and still have nothing to do with your choice.
Given what this benchmark actually measures: [summary], and what I am using it for: [your decision - picking a model, writing a policy, justifying a choice to someone],<br/><br/>tell me whether it is evidence for my decision, weak evidence, or irrelevant to it.<br/><br/>If it is irrelevant, say what evidence WOULD bear on my decision and whether it exists.<br/><br/>I would rather be told the benchmark everyone is quoting does not answer my question than make the decision on it.