MON, SEPTEMBER 28, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ General

Telling a Model It May Say Nothing Cuts Invented Data by Two Thirds

A web-extraction benchmark found models fabricated 70.7% of absent fields until one instruction - use null, do not guess - dropped it to 20.2%, with Gemini 3.8 Flash and GLM 5.3 at 2.8% and a paid API at 66.7%.

By AIToolsRecap September 28, 2026 8 min read 22 views
Home › Articles › General › One Sentence Cut Made-Up Fields From 71% to 20%

The finding

A benchmark set extraction models a deliberately unfair task: twin web pages differing by a single row, where one page contains the requested information and the other does not. The question was not whether models can extract - it was whether they admit when there is nothing to extract.

Without an instruction to abstain, models invented 70.7% of the missing fields.

With this added to the prompt, 20.2%:

Use null for any field whose value is not on the page. Do not guess.

Same models, same pages. One sentence, a 50-point swing.

The per-model results are the interesting part

  • Gemini 3.8 Flash - fabricated 1 of 36 missing fields (~2.8%)
  • GLM 5.3 - 1 of 36 (~2.8%)
  • Firecrawl, a paid extraction API - 24 of 36 (66.7%)

A free-tier flash model beat a paid extraction service by a factor of twenty-four on the one dimension that matters for extraction: not inventing data.

If you are paying for an extraction API, that is the number to take to your vendor. Speed and coverage are easy to demo. Abstention is not, and nobody advertises it.

Verification is cheaper than you think

The benchmark also tested whether a second cheap model could catch the fabrications. GPT-6 Luna caught 38 of 49 fabricated values while rejecting zero correct answers.

That second number is what makes this practical. A verifier that catches 78% of errors is useful; a verifier that catches 78% and never throws away a good answer is deployable, because the only cost of running it is the inference.

The architecture this implies: extract with whatever model fits your budget, then pass every non-null field to a cheap verifier asking only "is this value actually on the page". You are not paying for a better extractor, you are paying for a second opinion, and the second opinion is the cheap part.

What to change in your own prompts today

Add the abstention clause explicitly

Not "be accurate" and not "do not hallucinate" - those are instructions the model cannot act on. Give it a concrete thing to output when the answer is absent:

Use null for any field whose value is not on the page. Do not guess.

The mechanism is that you have given it a legal, non-failing answer for the missing case. Without one, returning nothing looks like failing the task, so it produces something plausible instead.

Make the schema allow absence

If your JSON schema marks every field required, you have instructed the model to fill them. Mark genuinely optional fields nullable and the abstention clause has somewhere to land.

Distinguish "not present" from "not found"

Ask for null when the field is absent from the source and a distinct marker when the model is unsure. Those are different failures and they deserve different handling downstream.

Verify before you store

One cheap model, one question - "is this value present in the source text" - applied to every non-null field. On these numbers it removes roughly three-quarters of your fabrications at a fraction of extraction cost.

The traps the benchmark used

Worth knowing because they are the ones that appear in real data: outdated prices still on the page, and misattributed authors. Both are cases where a plausible-looking wrong value is sitting right there, which is harder than a field being simply absent. 42 page pairs across 7 page types.

The honest caveats

  • Single run, synthetic pages. The pages were constructed for the test. Real pages are messier in ways that could push the numbers either direction.
  • Free tiers only. Paid API tiers were not tested beyond their free versions, so Firecrawl's paid tier may behave differently.
  • 36 missing fields per model is a small denominator. "1 of 36" and "2 of 36" are not meaningfully different results.

None of that undermines the headline. A 70.7% to 20.2% swing from one sentence is far too large to be noise, and it costs nothing to adopt.

FAQ

What exactly is the prompt?

"Use null for any field whose value is not on the page. Do not guess." Added to an otherwise unchanged extraction prompt.

Does this work outside web extraction?

The benchmark only tested extraction. The mechanism - giving the model a valid way to say nothing - is general, but the 50-point figure is specific to this task.

Which model was most honest?

Gemini 3.8 Flash and GLM 5.3, each inventing 1 of 36 missing fields.

Is a verification pass worth the cost?

On these results, yes - GPT-6 Luna caught 38 of 49 fabrications and rejected zero correct answers, so the only cost is inference on fields you already extracted.

Tags
AI GuideProductivity2026
⚑

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →
💡 AI Tools prompts
Prompt Guide
Best Claude AI Prompts for SEO (2026) — Content, Technical, and Comparison SEO
Claude Sonnet 5 and Opus 5 are strong for SEO work that requires writing quality, structured analysis, and long-form content generation. With 1M context, Claude can analyse an entire site's content structure, compare competing pages, and write complete article drafts in one session. These prompts cover the full SEO workflow: keyword research synthesis, content briefs, on-page optimisation, meta descriptions, technical audit interpretation, and comparison content that ranks above AI Overviews.
Get Prompts →
Prompt Guide
Best ChatGPT Prompts for SEO (2026) — GPT-5.6 and Browse
ChatGPT with GPT-5.6 Sol and Browse enabled is a capable SEO research tool — it can search the live web, analyse SERP results, and synthesise content briefs in a single session. GPT-5.6 Terra at $2.50/M offers a cost-efficient option for high-volume SEO content generation. These prompts are optimised for ChatGPT Plus with Browse, the ChatGPT Work product for larger projects, and the OpenAI API with web_search tool enabled.
Get Prompts →
Prompt Guide
Best Claude Opus 5 and Sonnet 5 Prompts for Writing (2026)
Claude Opus 5 and Sonnet 5 consistently produce the highest-quality long-form writing of any AI model in July 2026 — a lead documented across writing benchmarks and user testing since Claude 3 Opus. With 1M context and 128K output on Opus 5, Claude can write book chapters, complete reports, and long-form content without truncating. Sonnet 5 at $2/$10/M (intro through August 31) is the best value writing model available. These prompts are optimised for claude.ai Pro/Max, Claude Cowork, and the API.
Get Prompts →