AI NewsDev toolsAnnouncement

Answer me with HTML, a new agent skill, cuts the output tokens of a rendered answer by 7 times

An agent skill that renders answers as a one-page HTML took 923 output tokens where asking the model for HTML directly took 6,873, and finished in 13 seconds instead of 46.

AI News

Editorial2 min read

LinkedInX
GitHub social card for the answer-me-with-html repository

Image: GitHub

Why it mattersA reader waits for output tokens, so the lesson for a product team is to let the model write the content and let a template render the page.

Asking a coding agent to explain the TCP three-way handshake usually ends in a wall of text or, if you ask for HTML, a long wait for the model to type every div, class name and SVG coordinate. Answer me with HTML, a one-week-old agent skill that reached 234 stars on GitHub, keeps the HTML and asks the model for only the content.

The project is on GitHub as QingYunA/answer-me-with-html under MIT, created on 2 October. It installs as a Claude Code plugin, a global skill for Codex, Cursor and OpenCode, or by copying a folder, and once installed the model watches for questions whose answer is a diagram, a comparison table, a sequence or a timeline, and renders the page in about 50 milliseconds.

The measurement

The author ran the same questions through the same model (Claude Sonnet 5.5) both ways, three topics and three runs each, and reported the medians. Asking the model to write HTML directly used 6,873 output tokens, took 46 seconds and cost $0.22 per answer. Asking the content only, through the skill, used 923 output tokens, took 13 seconds and cost $0.26 per answer. That is 7.4 times fewer output tokens and 3.6 times less waiting, at roughly the same price.

The explanation given in the README is why the price does not fall with the tokens: the skill adds two short turns, one to load the skill and one to run the renderer, and every turn re-reads the conversation context. You save the wait, not the dollars. The per-topic numbers and the script that produced them live in the bench/ folder of the repository.

When the skill fires

The model decides whether a page is worth rendering. A one-line question like "how do I show hidden files with ls" gets a one-line answer. A multi-step flow, a comparison with pros and cons, or a request for a diagram triggers a page. An example-ask list in the README names six: a TCP handshake sequence diagram, a module call graph for a repo, a Redis-against-Memcached comparison with a verdict, per-sentence writing feedback with the problem words marked, a Kubernetes timeline, and the ls one-liner that stays as plain text.

Pages are saved to ~/.answer-me-with-html/pages/ and open in the browser unless the user turns that off. Themes and copy-the-source buttons are built in. There is an always-on mode behind a second plugin that appends a short page to every reply with a conclusion, costing about 90 tokens per turn in the reminder it adds.

For a product team

The useful idea here is older than the skill: when the answer is going to be rendered into a page, the model should write only the content, and a prepared layout should wrap it. A chat interface that returns a one-page explainer in 13 seconds will feel to a reader like a different product from one that takes 46, even when the content is the same. The skill demonstrates it on local agents and the method is listed step by step.

Source

Primary source: QingYunA/answer-me-with-html on GitHub.

SourceGitHub

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX