Models & agents

Kuber Mehta argues Minecraft-in-one-prompt and the pelican-on-a-bicycle SVG are demo benchmarks that labs plainly optimise for by the next launch

September 7, 2026 at 8:50 PM PT

Header image for Kuber Mehta's essay on demo benchmarks

Image: Kuber Studio

Why it mattersA team choosing a coding model on a launch-day demo is grading how much time the lab had to prepare, so a lead should buy on private evals or the team's own tasks rather than on the pelican.

Kuber Mehta published a short essay at kuber.studio on 6 September 2026, arguing that the viral one-prompt exercises that follow every model launch are marketing demos in the shape of tests. Each one is a fixed public target the lab has eight weeks to prepare for, so the next model is bound to look perfect on it. Hacker News surfaced the piece at 73 points as of the candidate sweep.

The tests being made in

Mehta names the five that filled his feed the hour GPT Astra shipped: recreating Minecraft in one prompt, painting the model itself in MS Paint, the pelican riding a bicycle rendered as an SVG, a ball bouncing in a rotating box with believable gravity, and an SVG game controller. His argument is that each one is visual, understandable and finite, which is what makes it launch material. Finite is also what makes it grade-able on a schedule. A test that never changes is a test that can be overfit, and a launch is the moment to show the work.

He grants that the labs are not being dishonest: nothing sells a launch like a pelican or a 3D game controller the timeline can quote. The problem is that the resulting number does not tell an engineer choosing a model whether the new one will be better at real work.

The parallel he draws to public leaderboards

Mehta extends the argument to the open scoreboards. He cites Thinking Machines' Inkling Small scoring within a point of its flagship sibling on the Artificial Analysis Intelligence Index with less than a third of the parameters, and beating it on Humanity's Last Exam, GPQA Diamond and SciCode. His framing is that public, static, famous test sets leak into training data and into fine-tuning choices, so smaller models that feel dumber in practice can still outscore better ones on those boards. Read at face value the leaderboard says the little one wins. Read carefully it says the little one was aimed at the leaderboard.

Mehta reports these numbers as claims from Thinking Machines and Artificial Analysis rather than as measurements he ran himself, and this write-up treats them the same way. The point of the essay lands on the shape of the incentives around any one score, rather than the score itself.

The alternative he lands on

The fix is a holdout eval. LiveBench rotates its questions, ARC-AGI keeps a private set, Humanity's Last Exam holds part of itself back. A model cannot be taught to a test that has not been written yet, so a private eval measures capability rather than preparation. Mehta accepts one caveat, and it is a good one: a private eval does not go viral. A demo benchmark makes it obvious to social media why a new capability is a step change, in seconds and without reading a paper. So the pelican still wins the launch, which is its own reward, and the useful move for anyone buying is to grade with something else.

For a technical lead reading the essay in the middle of choosing a coding model, the practical takeaway is familiar and now cleanly argued: the launch demo says the new model is good at the launch demo. If a model scores well on the round of quoted demos and keeps failing at the team's own tickets, that gap is the interesting number, and it is one the team can run in-house on a small private set of its actual tasks. A private eval is a quiet way to pick, and it survives the next launch cycle intact.

Source

Reported by: Kuber Studio

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

More from AI News

Latent Space published an AEO Tracker that measures how many sources each frontier model cites

Latent Space released a public tracker that runs 161 product categories through seven frontier models and records which sources each one cites, with median counts ranging from five for Astra to fifteen for Fable.

Source: PressGo-to-market

EEBench grades AI circuit designs with SPICE, and the best model scores 61.6%

EEBench published its September 1 leaderboard for AI-designed circuits, where Claude Opus 5 leads on 61.6% across 13 tasks graded by SPICE simulation rather than by a model judging the output.

Source: Hacker NewsModels & agents

Ai2 ran 16 benchmarks through item response theory and found the whole set collapses to two dimensions

The Allen Institute for AI trained a method called BenchMIRT on results from 100 models across 16 benchmarks and more than 34,000 questions, and reports that the whole set collapses to two underlying dimensions, with 10 percent of the questions preserving nearly the same picture.

Source: Vendor blogModels & agents