How to scale your engineering team / Keep quality high
How to measure engineering productivity
Most engineering productivity metrics measure activity, not progress. Lines of code, commits, tickets closed, and hours logged all reward looking busy. None of them tell you whether the product got better. As you scale, this matters more, because a bigger team produces more activity, and only outcome measures show real progress.
Published July 27, 2026. Editorial.
Key takeaways
- Output metrics like lines of code and tickets closed reward activity, not progress, and people learn to manipulate them the moment you track them.
- Measure outcomes: did the product improve, did users get what they needed, did the team deliver something that mattered.
- Team-level delivery measures like cycle time and reliability are more honest than individual output counts.
- Bad metrics do real harm at scale, because a growing team optimizes toward whatever you reward.
The urge to measure engineering productivity gets stronger as a team grows, and for good reason. With eight people you can feel whether the team is doing well. With forty you cannot, so you turn to numbers. The danger is that the easy numbers are the wrong ones, and the wrong numbers actively push the team in the wrong direction.
Why output metrics fail
Lines of code, commits, pull requests, story points, tickets closed, hours logged. These are all output metrics, and they share one basic flaw: they measure activity, not results. More lines of code is not better software, it is often worse. More tickets closed can mean the work was split into smaller tickets, not that more got done.
Output metrics also get manipulated as soon as you tie anything to them, and not because engineers are dishonest. People optimize for what is measured, so if you reward tickets closed, you get more, smaller tickets. If you reward lines of code, you get verbose code. The metric goes up and the product does not get better. We wrote about this directly in measure engineers by outcomes, not output, because it is one of the most common and most expensive mistakes in a scaling team.
Measure outcomes instead
The honest question is not how much did the team produce, but did the right things happen. Did the product improve in ways users can feel? Did the feature you released get used and solve the problem it was meant to solve? Did the team make progress on the goals that actually matter to the business?
Outcome measures are harder to collect and much harder to manipulate, which is exactly why they are worth the effort. An engineer cannot inflate an outcome the way they can inflate a commit count. And an outcome focus changes behavior in the right direction: it rewards the engineer who deletes a feature nobody used as much as the one who releases a new one, because both improved the product. This is the same principle that should drive how you structure engineering teams: give teams clear outcomes to own, then measure whether they achieve them.
Team measures beat individual counts
If you want numbers, and at scale you will, look at team-level delivery measures rather than individual output. How long does it take a change to go from started to released? How often does the team release? When something breaks, how fast is it fixed, and how often does a change cause a break in the first place?
These are the kind of measures Google's DORA program has studied across thousands of teams, and the useful finding is that the strongest teams release more often and cause fewer failures at the same time. Speed and stability rise together rather than trading off. That makes these measures a good main goal, because improving them almost always means the team got genuinely better, not just busier. Just keep them at the team level. The moment you rank individuals by them, you are back to manipulation.
What good measurement protects at scale
The reason this matters so much when you scale is that a team optimizes toward whatever you reward. With a small team, a bad metric does limited damage because everyone can see what is really happening anyway. With a large team, the metric becomes the only view leaders have, and the whole team starts working toward it. If that metric is tickets closed, you get a large team producing a lot of activity and not much product.
So the measurement you choose is not just a reporting decision. It decides the team's direction. Reward outcomes and a growing team keeps working toward real progress. Reward output and a growing team gets good at looking busy. This connects directly to the problems in why adding engineers can slow you down, because a team measured on output will happily grow headcount and slow down while every dashboard shows good results.
Keep it simple and honest
You do not need an elaborate system. Pick a small number of outcome and team-delivery measures, look at them over time rather than in single snapshots, and pair them with the judgment of senior people who can see what the numbers miss. Talk to the team about what the numbers mean instead of ranking people by them. Measurement at its best is a conversation about whether the work is succeeding, not a ranking. Do that and your metrics will still be telling you the truth when the team is three times its current size.
Common questions
What is the best way to measure engineering productivity?
Measure outcomes rather than output. Ask whether the product improved, whether released features got used and solved the problem, and whether the team made progress on goals that matter to the business. Outcome measures are harder to manipulate than counts like lines of code or tickets closed.
Why are lines of code and tickets closed bad metrics?
Because they measure activity, not results, and people manipulate them the moment you track them. Reward tickets closed and you get more, smaller tickets. Reward lines of code and you get verbose code. The number goes up while the product does not get better.
Should I measure individual engineer productivity?
Be careful. Individual output counts get manipulated and push the wrong behavior. If you want numbers, use team-level delivery measures like cycle time, release frequency, and how often changes cause breaks. Keep them at the team level, because ranking individuals by them recreates the manipulation problem.
What are good team-level engineering delivery measures?
Cycle time, how often the team releases, and when something breaks, how fast it is fixed and how often a change causes a break in the first place. Research on engineering teams associated with Google's DORA program has found the strongest teams release more often and cause fewer failures at the same time, so these measures move together rather than trading off.
Why is measuring engineering productivity harder as a team scales?
Because a bigger team produces more activity, which makes output metrics like lines of code or tickets closed look busier without the product actually improving. With eight people you can feel whether the team is doing well; with forty you cannot, so leaders turn to numbers, and the easy numbers available are usually the wrong ones to reward.
Can productivity metrics be manipulated by engineers?
Yes, almost automatically, once anything is tied to them. People optimize for what is measured, not because they are dishonest but because that is how incentives work. Reward tickets closed and you get more, smaller tickets. Reward lines of code and you get more verbose code, while the underlying product does not actually get better.
How do outcome-based metrics differ from output-based metrics?
Outcome metrics ask whether the product improved, whether a released feature got used, and whether the team progressed on goals that matter to the business. Output metrics count activity, like commits or tickets closed, regardless of whether it helped anyone. Outcome measures are harder to manipulate, because an engineer cannot inflate a real outcome the way they can inflate an activity count.
What is the risk of using bad productivity metrics at scale?
A growing team optimizes toward whatever you reward, so a bad metric does more damage as headcount rises. With a small team, everyone can see what is really happening regardless of the metric. With a large team, the metric becomes the only view leaders have, and the whole team works toward it, which means a metric like tickets closed can produce plenty of activity and little real progress.
Related reading
You cannot measure engineers by how much they produce
Lines of code, tickets closed, and hours logged all measure activity rather than progress. Here is how we think about engineering output without numbers that only look good.
How to set engineering OKRs that measure outcomes, not activity
"Release five features this quarter" is a to-do list, not a goal. Good engineering OKRs measure whether users and the business are better off, which is a much harder and more useful thing to write down.
More in Keep quality high
How to structure engineering teams
The structure of a team decides how much of its capacity turns into released product. As you grow, the connections between people multiply faster than the people, and every connection is a place work slows down. Good structure keeps most work inside small teams that own clear parts of the product, so growth adds speed instead of overhead.
How to onboard engineers quickly
Every engineer you add is a cost until they are productive, so how fast you onboard decides whether scaling pays off. The goal is real contribution in weeks, not months. That comes from three things: giving context instead of just a ticket, pairing new people with someone who knows the code, and writing down how the system actually works.