AI NewsOpen sourceAnnouncement

Game the LLM Reviewer rewrites papers to sit better with AI reviewers, without changing the science

A six-day-old MIT skill for Claude Code, Codex, Cursor and other coding agents applies six rhetorical strategies to a finished paper so an LLM reviewer scores it higher, while leaving every claim, number and citation untouched. The repository has 192 stars.

AI News

Editorial3 min read

LinkedInX
GitHub social card for the game-the-llm-reviewer repository

Image: GitHub

Why it mattersIf the person reading your writing is a language model, word order changes the score more than the science does, and the same problem is coming for landing pages, product descriptions and every piece of copy an AI ranks.

Move five words inside two sentences of a finished paper and an LLM reviewer's score goes from 6 out of 10 to 7. That is the working demonstration in Game the LLM Reviewer, an MIT-licensed skill Michael Jiahao Zhang published on GitHub on 21 September for Claude Code, Codex, Cursor, GitHub Copilot and other coding agents. It has 192 stars and 5 forks in six days.

The tool is a final editing pass a coding agent runs on a finished manuscript. The author writes and revises the paper as usual, then asks the agent to apply the skill. It produces a revised copy and a change note explaining every edit. The skill draws no reviewer model of its own and needs no reviewer API. Zhang states that the science stays the same and the author remains responsible for the manuscript and for their venue's disclosure rules.

What actually changes

The skill defines six operations, five that edit and one that checks. Contribution stance moves an established property into the sentence that opens the method description. Evidence framing rephrases the same numerical result. Abstract emphasis reorders existing statements without adding new ones. Lexical stance adjusts non-factual wording while preserving certainty. Scope framing rewords the evaluated and unevaluated cases without reducing a limitation. A sixth step, equivalence check, tests each edit against the original.

Zhang publishes three worked examples on real, published papers. In the ToolLLM abstract, the phrase "to evaluate ..., we develop an automatic evaluator: ToolEval" was reordered to "we develop ToolEval, an automatic evaluator, to evaluate ..." and the score moved from 6 to 7 out of 10. In the API-Bank abstract, "Lynx surpasses Alpaca's tool utilization performance by more than 26 pts" became "relative to Alpaca, Lynx improves tool utilization performance by more than 26 pts", and the score also moved from 6 to 7. In the WebArena introduction, "We focus on evaluating the functional correctness" became "Our evaluation focuses on the functional correctness", and the score moved from 7 to 8. Zhang says every other sentence stayed the same, and states that the scores are single independent readings from GPT-6 ASTRA on a 10-point scale, from one rewriter pass with no reviewer feedback in the loop.

What the readme is careful about

The repository carries an academic integrity notice at the top. It asks that edits preserve claims, citations, assumptions, stated uncertainty and any substantive limitations, and that each change be recorded. Zhang writes that fabricated results, inflated novelty, hidden weaknesses and instructions to reviewers are outside the skill's scope, and that authors remain responsible for the manuscript. He also opposes handing peer-review decisions to LLMs, and describes the tool as defensive polishing for authors who have no say in whether a reviewer uses one.

The venue rule matters. Some conferences and journals now require disclosure of any AI assistance in writing, and some ban it. A paper edited by this skill without the required disclosure is a rule break at those venues, whatever the content.

The narrower reader interest here is academic writing under LLM review. The broader point applies to anyone whose text is judged by a language model, from a landing page an answer engine will cite, to a product description a shopping assistant ranks, to a bio a hiring model scores. Word order changes the score more than the substance does.

Source

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX