The finding
A benchmark set extraction models a deliberately unfair task: twin web pages differing by a single row, where one page contains the requested information and the other does not. The question was not whether models can extract - it was whether they admit when there is nothing to extract.
Without an instruction to abstain, models invented 70.7% of the missing fields.
With this added to the prompt, 20.2%:
Use null for any field whose value is not on the page. Do not guess.
Same models, same pages. One sentence, a 50-point swing.
The per-model results are the interesting part
- Gemini 3.8 Flash - fabricated 1 of 36 missing fields (~2.8%)
- GLM 5.3 - 1 of 36 (~2.8%)
- Firecrawl, a paid extraction API - 24 of 36 (66.7%)
A free-tier flash model beat a paid extraction service by a factor of twenty-four on the one dimension that matters for extraction: not inventing data.
If you are paying for an extraction API, that is the number to take to your vendor. Speed and coverage are easy to demo. Abstention is not, and nobody advertises it.
Verification is cheaper than you think
The benchmark also tested whether a second cheap model could catch the fabrications. GPT-6 Luna caught 38 of 49 fabricated values while rejecting zero correct answers.
That second number is what makes this practical. A verifier that catches 78% of errors is useful; a verifier that catches 78% and never throws away a good answer is deployable, because the only cost of running it is the inference.
The architecture this implies: extract with whatever model fits your budget, then pass every non-null field to a cheap verifier asking only "is this value actually on the page". You are not paying for a better extractor, you are paying for a second opinion, and the second opinion is the cheap part.
What to change in your own prompts today
Add the abstention clause explicitly
Not "be accurate" and not "do not hallucinate" - those are instructions the model cannot act on. Give it a concrete thing to output when the answer is absent:
Use null for any field whose value is not on the page. Do not guess.
The mechanism is that you have given it a legal, non-failing answer for the missing case. Without one, returning nothing looks like failing the task, so it produces something plausible instead.
Make the schema allow absence
If your JSON schema marks every field required, you have instructed the model to fill them. Mark genuinely optional fields nullable and the abstention clause has somewhere to land.
Distinguish "not present" from "not found"
Ask for null when the field is absent from the source and a distinct marker when the model is unsure. Those are different failures and they deserve different handling downstream.
Verify before you store
One cheap model, one question - "is this value present in the source text" - applied to every non-null field. On these numbers it removes roughly three-quarters of your fabrications at a fraction of extraction cost.
The traps the benchmark used
Worth knowing because they are the ones that appear in real data: outdated prices still on the page, and misattributed authors. Both are cases where a plausible-looking wrong value is sitting right there, which is harder than a field being simply absent. 42 page pairs across 7 page types.
The honest caveats
- Single run, synthetic pages. The pages were constructed for the test. Real pages are messier in ways that could push the numbers either direction.
- Free tiers only. Paid API tiers were not tested beyond their free versions, so Firecrawl's paid tier may behave differently.
- 36 missing fields per model is a small denominator. "1 of 36" and "2 of 36" are not meaningfully different results.
None of that undermines the headline. A 70.7% to 20.2% swing from one sentence is far too large to be noise, and it costs nothing to adopt.
FAQ
What exactly is the prompt?
"Use null for any field whose value is not on the page. Do not guess." Added to an otherwise unchanged extraction prompt.
Does this work outside web extraction?
The benchmark only tested extraction. The mechanism - giving the model a valid way to say nothing - is general, but the 50-point figure is specific to this task.
Which model was most honest?
Gemini 3.8 Flash and GLM 5.3, each inventing 1 of 36 missing fields.
Is a verification pass worth the cost?
On these results, yes - GPT-6 Luna caught 38 of 49 fabrications and rejected zero correct answers, so the only cost is inference on fields you already extracted.