Here is a benchmark, eval or study I am about to cite: [paste the paper, repo README, methodology page or leaderboard notes].<br/><br/>Find and quote exactly:<br/>- what the authors state it measures<br/>- what they state it does NOT measure<br/>- every limitation or caveat they put in writing themselves<br/><br/>Put the authors' own disclaimers first, before any summary of results. If they describe it as exploratory, preliminary or unvalidated, lead with that.<br/><br/>If there is no methodology documentation at all, say so - that is the finding.
For this benchmark: [paste], tell me precisely what the scores are measured AGAINST.<br/><br/>Is it a human baseline, and if so which humans, sampled how, when? Is it another model? An absolute standard? A survey average?<br/><br/>Then tell me what "scoring well" literally means in those terms - not what it sounds like it means.<br/><br/>If the reference is a population average, say plainly that closer to average is what is being rewarded, and what that does and does not imply.
This evaluation adapts an existing instrument, test or dataset: [describe or paste].<br/><br/>Tell me what was changed - items replaced, wording altered, scales rescored, prompts restructured - and what the original validation covered.<br/><br/>Then tell me honestly whether the original validation still applies after those changes, and which specific claims it no longer supports.<br/><br/>Do not be generous about this. A validated instrument with altered items is a new instrument.
Here are a benchmark's results: [paste].<br/><br/>Before interpreting any rank, tell me:<br/>- how many runs per model<br/>- what the variance or self-consistency across runs was<br/>- whether error bars, confidence intervals or agreement rates are reported at all<br/><br/>Then tell me which differences between adjacent ranks are smaller than the run-to-run variation, and therefore are not differences.<br/><br/>If no reliability figure is reported, say that the ranking is one sample and should be read as such.
Here is how this result is being reported: [paste the headline, post or summary].<br/>Here is what the benchmark actually measures: [paste from the earlier prompts].<br/><br/>Rewrite the claim so it is accurate and still worth saying. Keep it short.<br/><br/>Then tell me which parts of the original are overstatement, which are category errors, and which are simply wrong.<br/><br/>If the accurate version is boring, say so - that tells me whether the result was worth citing at all.
Given what this benchmark actually measures: [summary], and what I am using it for: [your decision - picking a model, writing a policy, justifying a choice to someone],<br/><br/>tell me whether it is evidence for my decision, weak evidence, or irrelevant to it.<br/><br/>If it is irrelevant, say what evidence WOULD bear on my decision and whether it exists.<br/><br/>I would rather be told the benchmark everyone is quoting does not answer my question than make the decision on it.