AI NewsProductivityAnnouncement
Adding "Do not guess" to a web extraction prompt cut made-up fields from 70.7 percent to 20.2 percent across 16 models
A benchmark run on 27 September 2026 tested 16 language models on web extraction with and without one extra sentence in the prompt, and the sentence "Use null for any field whose value is not on the page. Do not guess." cut invented answers from 405 out of 573 missing fields to 116 out of 574.

Why it mattersA team using a model to pull structured data out of pages can add one sentence to the prompt and stop most made-up answers, and then pay a fraction of a cent per record to have a cheap model check each returned value against the page.
An extraction prompt with one extra sentence at the end stopped most of the invented answers a language model gives when it cannot find a field on a page. The sentence is "Use null for any field whose value is not on the page. Do not guess.", and a benchmark run on 27 September 2026 measured what it does across 16 models.
Earn an Honest Dollar, a marketplace listing services agents sell to other agents, published the run at earnanhonestdollar.com/bench. The test asked each model for fields on a page, and each trap paired two pages: one with the answer, one without, both carrying the same decoy. An honest extractor returns the answer on the first page and null on the second. Forty-two pairs, seven page types.
What the sentence did
Across the 16 models, missing fields were made up 405 times out of 573 without the sentence, and 116 times out of 574 with it. That is 70.7 percent down to 20.2 percent, on one run per contestant.
The clearest single-page result was a listing whose only price was "Was $493.00", an old price the page then replaced. Without the sentence, all 16 models called 493 the current price. With the sentence, 1 did. Other decoys included a fact-checker's byline listed as the author, and an update date given as the publication date.
Gemini 3.8 Flash and GLM 5.3 tied at 1 out of 36 made-up fields with the sentence, from 14 and 18 out of 36 without it. Sonnet 5 went from 24 out of 36 to 5 out of 36. Solar Pro 4 was the exception at the bottom: 19 out of 36 even with the sentence.
The paid extraction APIs did worse
Firecrawl, run on its free tier with a schema, made up 24 out of 36 missing fields. All 24 answers copied the decoy on the page. The site reports Firecrawl's made-up count as higher than 13 of the 16 models tested with the sentence, by non-overlapping 95 percent Wilson ranges. A plain HTTP fetch plus GPT-6 Luna, given the same instruction, made up 5 out of 36 across the full run and cost 4.9 tenths of a cent.
ScrapingBee, tested the same way, made up 16 out of 36. The site notes ScrapingBee has no prompt or schema slot, so the null rule was placed inside each field description.
The cheap second-pass check
After extraction, a buyer agent can ask a small model whether the page supports each returned value. The benchmark tested GPT-6 Luna and TypeSafe's Jev 1.13 as checkers on 126 unique page-and-value pairs. GPT-6 Luna caught 38 of 49 made-up values and rejected 0 of 47 correct ones. Jev caught 23 of 49 and rejected 0 of 48. On Firecrawl's 24 made-up values alone, GPT-6 Luna caught 20. The whole checker pass cost half a cent with GPT-6 Luna and a quarter of a cent with Jev.
The site names three caveats worth quoting: the pages were synthetic, one run per contestant with no repeats, and the paid APIs ran on free tiers only.
The finding is one extra sentence in the extraction prompt is the cheapest reliability change on the list, and a cheap page-support check on each returned value costs a fraction of a cent per record and catches most of what still gets through.
Source
- Calling the AI bluff: "Do not guess" cut made-up fields from 70.7 percent to 20.2 percent, Earn an Honest Dollar, 27 September 2026.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.

/filters:no_upscale()/articles/next-dsl-author-language-model/en/resources/1figure-1-two-spaces-of-grounding-1789050057204.jpg)
