AI NewsProductivityReported

Claude Code's pre-filled suggestions could be training data, an engineer argues

Zohaib Ansari argues on his blog that Claude Code's new pre-filled next-message feature is a cheap way to collect preference pairs from working engineers, with every edited suggestion a labelled correction from someone who knows the project.

AI News

Editorial3 min read

LinkedInX
Zohaib Ansari's blog post about the Claude Code suggested message feature

Why it mattersA team shipping an AI product that needs preference data has a design pattern here, prompt the user with your best guess and treat every edit as a labelled correction, instead of paying annotators who do not know the codebase.

Claude Code's prompt box now fills in a suggested next message after a turn, something like "run the tests" or "commit this". The user can send it as it is, edit it, or ignore it. Zohaib Ansari, a software engineer in Waterloo, argues in a blog post that this convenience saves a few seconds for the user and does something more useful for Anthropic, which is collecting preference data from working engineers at the moment they most care about getting it right.

The feedback problem the suggestion box solves

Every AI product has thumbs-up and thumbs-down buttons, and few users click them. The ones who do tend to be annoyed, Ansari writes, so the labels are sparse and biased. Paid annotators are expensive and an annotator reading someone else's codebase guesses at what the developer wanted. A suggestion sitting in the prompt box gets around both problems. Each one is a prediction of the user's next turn, conditioned on the whole session up to that point. The user grades it without thinking of it as grading.

Sending a suggestion unchanged counts as a positive label. Editing it is worth more. The original and the edited version form a preference pair, and the difference between them shows where the prediction went wrong. Ansari gives the example of a suggestion that reads "run the tests" and gets changed to "run only the auth tests, the full suite takes ten minutes". That is a correction from someone who knows the project, written at the moment they care about getting it right.

The second training objective

Ansari also notes a benefit the suggestion box carries on its own, before any preference modelling. Predicting what a competent developer asks for next is a useful training objective for an agent, because it rewards the model for learning how work is sequenced. After a refactor you run the tests. Once they pass you commit. That is a short step from an agent that takes the next action without being asked.

The post describes the shape of the data a classic RLHF pipeline needs: human preferences between model outputs, used to train a reward model that then shapes the policy. The expensive part has always been collecting the preferences. The suggestion box collects them from people doing their jobs, on real codebases rather than a synthetic benchmark, so the labels are in distribution for the work the model will later be asked to do.

What Ansari does not claim

Ansari is explicit that this is his reading rather than inside knowledge. "I have no inside knowledge of what Anthropic does with this signal, and whether your sessions are used for training at all depends on your plan and privacy settings." He adds that if he had to design a way to collect preference data from thousands of working engineers, he would struggle to come up with something cheaper. It costs one line of text in an input box and ships as a convenience.

The pattern generalises past Claude Code. Any AI product with a text input can ship its best guess for the next turn, let the user send, edit or ignore it, and keep the three signals. On Hacker News the post reached 255 points on the day it ran.

Source

The smartest Claude Code feature, Zohaib Ansari.

Reported byZohaib Ansari

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX