You cannot measure engineers by how much they produce

A manager once told us, with real pride, that his most valuable engineer had written the most code that quarter. We asked what problems that code had solved. He did not have an answer, because he had never been measuring that. He had been measuring volume, because volume is easy to count, and easy to count is not the same as worth counting. Engineering is full of metrics that measure activity and call it progress. Here is how we try to avoid that mistake.
1. Output volume rewards the wrong behavior
The metrics that get used most are the ones that are easy to pull: lines of code, tickets closed, hours logged, commits made. They share a serious flaw. They all measure how much someone did, not whether what they did mattered. And the moment you reward a number, people optimize for that number, so you get exactly what you asked for and not what you wanted. This is old and well understood. It is Goodhart's Law: when a measure becomes a target, it stops being a good measure.
Reward lines of code and you get more code, including a lot that should not exist, because the simplest solution rarely wins when volume is the goal. Reward tickets closed and you get many small tickets and a reluctance to take on the hard, slow problem that would count as one ticket and take a month. The metric misses the value, and it also pushes people away from the most valuable work, which is often quiet and does not raise the count.
We saw a version of this at one company we worked with. Their weekly report ranked engineers by pull requests merged. Within a month the best people had learned to split one clean change into five small ones, each with its own review and its own approval, so the chart looked busy. The work took longer, the reviews got less careful, and the number went up. Nobody was being lazy or dishonest. They were doing what the number asked. That is the whole problem with a bad metric: it does not need bad people to produce bad outcomes.
We have felt this ourselves. Any time we let a volume number become the goal, behavior changed to fit the number and moved away from the real goal.
2. Measure the problem solved, not the code written
The honest alternative is harder, and that is exactly why most teams avoid it. Instead of asking how much someone produced, ask what problem got solved. Did the product get better, more reliable, or easier for the next person to work in? Did a whole class of bugs disappear? Did a confusing part of the system become simple? Those are outcomes, and outcomes are what you actually care about.
The difficulty is that outcomes are hard to put in a dashboard. You cannot automatically count "made the checkout flow trustworthy" the way you can count commits. It takes a manager who understands the work well enough to judge it, which is more effort than reading a chart. But the effort is the job. When you judge work by the problem it solved, people start solving problems instead of generating activity, and the whole team quietly gets more serious about what matters.
The strongest research agrees. The DORA / Accelerate State of DevOps research, which studies thousands of teams over many years, measures delivery by outcomes such as how fast a team can release a change and how stable that change is in production, and it treats output counts like lines of code as poor measures of value. The most credible research on the subject does not tell you to count more. It tells you to count the result.
The point is to measure the thing you want, even when it is harder to measure, rather than the thing that is easy and wrong.
3. The best engineers often delete more than they add
Here is the clearest sign that volume metrics are broken. Some of the most valuable work an engineer does is remove code, not write it. Every line in a system is a small ongoing cost. It has to be understood, maintained, and worked around by everyone who comes after. A senior engineer who deletes a thousand lines while keeping the same behavior has made the whole team faster for years, and on a line-count metric they just had a quarter with a large negative number.
The same is true of the features that never get built. Talking a team out of a complicated feature that would have slowed the team down and served almost no one is high-value work. It produces nothing you can count. It might be the best thing anyone did all month. If your way of valuing engineers cannot see that work, it cannot see some of the most important contributions on the team, and the people doing that work will eventually notice and stop.
We saw this happen on one project. An engineer spent most of a week not writing a feature. She read the request, talked to the people who asked for it, and found that a setting they already had would do the job with one small change. She released the small change and closed the larger request. On any volume metric that was one of her worst weeks all year. In reality she saved months of building and maintaining a feature nobody needed. The right question was what she prevented.
4. The number tells the team what you actually value
There is a second cost to a bad metric, and it is slower and larger than the first. A metric is a message. It tells everyone on the team, in plain terms, what the company rewards. People read that message and adjust, even the ones who know better, because over time it decides who gets promoted and who does not.
So the danger goes beyond manipulated numbers this quarter: a year of rewarding volume teaches your strongest people to stop doing the quiet, high-value work, because they can see it does not count. Under Goodhart's Law the target changes more than one report: it changes what the whole team believes good work looks like. This is why the DORA research treats a small set of outcome measures as a shared measure for the team rather than a ranking of individuals. A good measure directs everyone to the same real goal. A bad one pulls the team's effort away from it and calls the result performance.
The fix is to say out loud what you count as valuable and then make sure your metrics agree with it. If you tell people that simplifying the system matters and then reward only the people who add to it, they will believe the reward, not the speech. Make the two agree, and you no longer have to convince anyone to do the right work. The way you measure will do it for you.
What this means for you
Look at how your team judges engineering work and ask what behavior it actually rewards. If the answer is volume, expect volume, including the kind you do not want. Shift the question from how much someone produced to what problem they solved, even though that takes a manager who understands the work. And make sure your idea of a valuable contribution has room for the engineer who simplified the system or prevented the wrong feature, because that person is often doing the most important work on the team.
Thanks to the engineers here who have deleted more than they have added and made everything better for it, and who kept doing the uncountable work anyway. Activity is easy to measure and easy to fake. Progress is neither, which is exactly why it is the only thing worth measuring.
Volume used to be a bad measure. Now it means nothing. When a model writes the implementation, an engineer can produce two thousand lines before lunch, and that number says nothing about whether the right thing was built or whether anyone checked it. Every output metric that was kept because it partly followed effort has lost even that reason. What is left is what should have been measured all along: did the problem go away, and does the system still behave correctly a month later.
References
- Goodhart's Law: Marilyn Strathern's 1997 wording, "When a measure becomes a target, it ceases to be a good measure." The reason rewarding a number changes behavior until the number stops meaning anything.
- DORA / Accelerate State of DevOps research: the field's most credible delivery research, which measures teams by outcomes such as throughput and stability and treats output counts like lines of code as poor measures of value.
Related guide: How to scale your engineering team.


