AI NewsModels & agentsAnnouncement
Opper's Pac-Man benchmark ran six decision models through 100 games each
Opper released an open-source Pac-Man benchmark that ran six System One decision models (jev 1.13, GPT-6 Luna, Clef, Clef Flash, Kev 4B and Laya) through 100 real-time games each against the classic arcade ghosts.

Image: Opper
Why it mattersA team picking a decision model for an agent now has a shared task that measures real-time play, cost per game and fallback rate side by side, so the choice stops being a vendor sales deck.
Every decision-model vendor claims a strong benchmark on its own test set. An open-source Show HN now has six of them playing the same arcade game.
The project is called jevman and runs each model as the brain steering Pac-Man against the classic arcade ghosts, 100 games per model. Opper released it on Hacker News, where the post reached 60 points. The code is AGPL-3.0 at opper-ai/jevman-benchmark and started as a fork of joch/jevman.
How the test runs
Each model plays 100 games against the scripted "classic" ghosts, with three lives per game and a five-minute cap on any single run. At every junction the game sends the maze, pellets, ghosts, fruit and distances to the model as JSON and asks for a direction. An answer must come back within two seconds, or a greedy backup rule takes over for that one decision and the game continues. The leaderboard is published at opper.ai/jevman-benchmark/leaderboard, scores are reported as the mean with a 95 percent margin of error, and the raw JSON is committed to the repository so any run can be replayed.
The leaderboard
TypeSafe's jev 1.13 led with a mean score of 2,750 (±218 at two standard errors) and the top single-game high score of 6,380, Opper says. OpenAI's GPT-6 Luna Decisions came second at 2,568, Cloudflare's Clef Flash third at 2,538, and Cloudflare's full-size Clef fourth at 2,476. Opper's own Kev 4B, a Qwen3.5-4B fine-tune, scored 1,506, and ConvAI Innovations' Laya scored 639, which Opper attributes to its 512-token context being shorter than one of jevman's questions.
Three numbers beyond score shift the picture. Clef Flash costs $0.0036 a game, Clef costs $0.0173, so the Flash variant scored higher at a fifth of the price. GPT-6 Luna had the lowest latency of the top three at 179 milliseconds versus jev's 290 and Clef's 398. Jev had the lowest fallback rate at 3.4 percent, meaning the greedy backup rule took over on fewer decisions than for Luna (4.4 percent) or Clef (4.6 percent).
Submitting a new model
Opper notes that any model behind an HTTP endpoint can enter the leaderboard as self-reported: run the harness against the endpoint, send the recorded games as a pull request, and CI replays each one and checks its score. The repo ships an example endpoint of 34 lines to start from. The hosted page also lets a visitor play Pac-Man against the models as the four ghosts, on either Opper's free credit pool or a signed-in Opper account, at about one or two cents a game.
A team building anything that calls a decision model many times a second now has a shared Pac-Man score to compare vendors against, instead of each lab's own benchmark page. The usual caveats apply: jevman tests only real-time play against one set of scripted ghosts, Opper serves four of the six models it compares, and Opper labels its own posture plainly in the README.
Source
Primary source: jevman: a Pac-Man benchmark for decision models, by Opper. Repository: opper-ai/jevman-benchmark, AGPL-3.0. Discussion: Show HN.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.

