MON, SEPTEMBER 28, 2026
Independent · In‑Depth · Practitioner‑Tested
Claude Productivity

Seven Extraction Prompts That Return Null Instead of Fiction

A benchmark this week found extraction models invented 70.7% of fields that were simply absent from the page. Adding one sentence - use null for anything not on the page, do not guess - took it to 20.2%. Gemini 3.8 Flash and GLM 5.3 got to 2.8%; a paid extraction API scored 66.7%. These seven prompts build that behaviour into the whole pipeline, from the schema to the verification pass.

⌨️ 7 prompts 🕐 Updated Sep 28, 2026
💡 How to use these prompts: Replace everything in [BRACKETS] with your specific details before sending. Click Copy to copy any prompt to your clipboard instantly.
1
The Base Extraction Prompt
Rule 1 is the sentence that moved the benchmark from 70.7% to 20.2%. Rules 2-4 close the ways a model reaches a value indirectly.
Extract the following fields from the page content below.<br/><br/>FIELDS:<br/>[list your fields]<br/><br/>RULES:<br/>1. Use null for any field whose value is not on the page. Do not guess.<br/>2. Do not infer a value from context, related fields, or general knowledge. If the exact value is not written on the page, it is null.<br/>3. Copy values verbatim. Do not normalise, reformat or correct them.<br/>4. If a field appears more than once with different values, return all of them as an array and do not choose.<br/><br/>Return JSON only.<br/><br/>PAGE:<br/>[paste]
2
Make the Schema Allow Absence
An all-required schema tells the model to fill every slot. The abstention instruction has nowhere to land.
Here is my extraction schema.<br/><br/>Rewrite it so that every field which could legitimately be missing from a source is explicitly nullable, and only fields that are genuinely always present are required.<br/><br/>For each field you change, say why it can be absent in real sources.<br/><br/>Then tell me which fields I have marked required that probably should not be - a required field is an instruction to the model to produce something, which is exactly what I am trying to avoid.<br/><br/>SCHEMA:<br/>[paste]
3
The Verification Pass
In the benchmark this pattern caught 38 of 49 fabrications while rejecting zero correct answers. Run it with the cheapest model you have.
Below is a source document and a set of field values extracted from it.<br/><br/>For each value, answer one question only: does this exact value appear in the source text?<br/><br/>Return: field name, the value, PRESENT or ABSENT, and if present the surrounding sentence.<br/><br/>Do not judge whether the value is correct, plausible or well-formatted. Only whether it is there.<br/><br/>SOURCE:<br/>[paste]<br/><br/>EXTRACTED:<br/>[paste]
4
Separate Absent From Uncertain
"Not there" and "I am unsure" need different handling downstream. Collapsing them into null loses the distinction.
Extract the fields below. For each one return an object with two keys: value and status.<br/><br/>status must be exactly one of:<br/>- FOUND - the value is written on the page<br/>- ABSENT - the page does not contain this information<br/>- AMBIGUOUS - the page contains something that might be it but you are not confident<br/><br/>For AMBIGUOUS, put your best candidate in value and explain the ambiguity in a third key, note.<br/><br/>Never use FOUND for a value you inferred.<br/><br/>FIELDS:<br/>[list]<br/><br/>PAGE:<br/>[paste]
5
Catch the Stale-Value Trap
Outdated prices and misattributed authors were the benchmark's hardest traps - a plausible wrong value sitting in plain sight is harder than an absent one.
The page below may contain outdated versions of the values I want - an old price still shown beside a new one, a previous author, a superseded date.<br/><br/>For each field: return the CURRENT value, and separately list any other candidate values you found with the reason you rejected each.<br/><br/>If you cannot tell which is current, return status AMBIGUOUS and list all candidates rather than picking.<br/><br/>FIELDS:<br/>[list]<br/><br/>PAGE:<br/>[paste]
6
Audit an Existing Pipeline
Few-shot examples where every field is populated teach the model that populated is the expected shape.
Here is my current extraction prompt and a sample of its output.<br/><br/>Identify every place the prompt implicitly pressures the model to produce a value: required fields, examples that are always fully populated, phrasing like "extract the price" that presupposes a price exists, and output formats with no representation for absence.<br/><br/>Rewrite the prompt to remove that pressure without losing extraction quality.<br/><br/>Then predict which of my sample outputs are fabricated and say what made you suspect each.<br/><br/>PROMPT AND SAMPLES:<br/>[paste]
7
Measure Your Own Fabrication Rate
The published benchmark used 42 page pairs across 7 page types. Your own numbers will differ; measure rather than assume.
Design a test I can run this week to measure how often my extraction pipeline invents data.<br/><br/>Specify: how to construct source documents where I know a field is genuinely absent, how many samples for a usable signal, what to record, and how to calculate the fabrication rate.<br/><br/>Include the twin-document technique - two near-identical pages differing only in whether the target value is present - and tell me what confounds it.<br/><br/>Give me a result table format I can fill in.<br/><br/>MY PIPELINE:<br/>[paste]