AI NewsProductivityAnnouncement
Work-humanizer rewrites AI drafts and keeps every commitment unchanged
Work-humanizer is an open-source agent skill that rewrites AI-written work text in a plain voice while keeping every commitment at the same strength, and in a blind test on 12 texts four AI judges picked its version 48 percent of the time against 20 percent for the next tool.
Image: work-humanizer
Why it mattersThe usual AI rewrite changes what the sentence commits to, which turns a sales email or a support reply into a different promise, so a rewrite skill measured against that failure mode is worth the test on your own drafts.
A sales email, a support reply, a legal letter: the exact words are a commitment, and the common fix for an AI-sounding draft rewrites around those commitments and quietly changes them. phuryn/work-humanizer, open-sourced under MIT on 9 October 2026 and 83 stars by the next afternoon, is an agent skill built to rewrite the AI tone out without touching what the text says.
The skill loads into Claude Code, Codex, Cursor or any agent that reads a SKILL.md, and the author published its evidence in the same repository. On 12 new AI-written texts, every tool ran on Claude Sonnet in a fresh session with the same wrapper and only the skill text changed. Four judges, Grok, DeepSeek, Qwen and OpenAI Codex, saw each candidate without labels in two random orders. Asked which rewrite they would send, with the writer's intent shown, they picked work-humanizer 48 percent of the time, blader/humanizer 20 percent, a one-line "rewrite this to sound human" prompt 1 percent, and Emulate-1 0 percent.
Why the choice gap exists
A fact checker read every output and listed each dropped, changed or added claim by hand. Work-humanizer made 18 meaning changes across the 12 texts. The other tools made 33, 42 and 120 in the same test. Across the full set of 56 texts the skill rewrote in all test configurations, it invented zero facts.
A one-line "sound more human" prompt beat the AI draft on 90 percent of judgments against work-humanizer's 75 percent, because it rewrites freely and adds things the writer never said. Judges who saw the original would send its version once in 96 picks.
No rewrite gets past Pangram, except with the writer's own reasoning
Pangram 4, a commercial AI detector, scored every output. Work-humanizer, blader/humanizer and the plain-prompt rewrite all left the text at 2 to 4 percent human on the 16 main test texts, where the AI draft started at 4 percent human. Emulate-1, which is sold as undetectable, reached 87 percent by changing the most. On four texts where work-humanizer was given the writer's own notes to work from, Pangram scored its version 100 percent human. The author concludes that word choice matters least and that handling the writer's own reasoning carries the detector score.
Interview mode asks before it writes
Work-humanizer carries a second mode for drafts missing information only the writer has. It asks up to five short questions and builds the piece from the answers. In a 16-text run where a simulated writer answered from a list of facts they knew, it asked on nine, used 39 of the writer's facts and invented none; judges would send its version 69 percent of the time against 22 for blader/humanizer. The simulation is not a real person test, which the author names in the caveats.
The repository ships a mechanical check script, check.py, that scans for the usual AI tells: em dashes rerouted into semicolons, uniform sentence lengths, "it's not X, it's Y" contrasts and stock vocabulary.
Caveats the author names first
The author lists the limits before the numbers. Samples are small: 12 texts on the scoreboard, 56 across all sets, four for the "with your notes" result. The texts are sales, marketing, support, product and internal writing, with no contracts or legal documents. The judges are AI, they disagree on what "human" means, and the Pangram runs are one version on one day. The method and raw data sit under evidence/ in the repository for anyone to read before quoting a figure.
Source
phuryn/work-humanizer on GitHub. Numbers for judge picks, meaning-change counts, Pangram scores, test method, caveats and the comparison tools used are all from the project's README and evidence/ folder. Star count and licence are from the GitHub API.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.