How to set engineering OKRs that measure outcomes, not activity

Every quarter a version of the same document gets written on engineering teams around the world. It lists the features the team plans to release, presents them as goals, and calls itself the OKRs. It looks responsible. It is easy to fill in. And it measures the one thing that matters least: how busy the team was. You can hit every item on it and leave your users exactly as frustrated as they were before, and your business exactly where it started. We have seen teams do precisely that, work through a full quarter of released features, and end up wondering why nothing improved. The problem was never the effort. It was that they measured the effort instead of the result.
Good engineering OKRs are harder to write, and that difficulty is the point.
Objectives and key results, and why the wording matters
A quick definition first, because the term is often used loosely. OKR stands for objectives and key results. The objective is the goal, written in plain words: the thing you are trying to make true. The key results are the few measurable signals that tell you honestly whether you got there. The whole idea is to name what you want to achieve and how you will know if you did.
The problem is in the key results. Most teams write key results that describe the work they will do, and the whole method stops working without anyone noticing. "Release the new search feature" is a key result you can hit by releasing something called a search feature, regardless of whether it helps anyone find anything. It is a to-do list item. The fix is to write the key result as the change you want in the world instead: "users find the right result on their first search 80 percent of the time." Now the work can succeed or fail honestly. You can release the feature and still miss the goal, which is exactly the feedback you need, because it tells you the work did not actually solve the problem.
Output is not outcome
The distinction behind all of this is between output and outcome, and it is worth being precise about it because everything depends on it.
Output is what you produce. Features released, tickets closed, code written, releases made. It is easy to see and easy to count, which is exactly why teams love to measure it. Outcome is the change that results from the output. Users completing a task they used to abandon, a page loading fast enough that people stop leaving, costs dropping, revenue rising. Outcome is harder to see and harder to count, and it is the only thing that actually matters.
The reason this confuses so many teams is that output feels like outcome in the moment. Releasing a feature feels like progress. It looks like progress on a status update. But a released feature that nobody uses, or that does not change the number it was meant to change, is not progress. It is activity. We wrote about this same confusion at the level of individual engineers in measure engineers by outcomes not output, and it applies just as much to a whole team's quarterly goals. If your key results count outputs, you are guaranteeing that a busy quarter will look like a successful one whether or not it was.
Tie engineering goals to what users and the business need
So where do good engineering OKRs come from? They come from working backward. Start with what a user or the business actually needs, then find the engineering work that changes it.
Say the business needs more of the people who start signing up to actually finish. That is a real outcome someone outside engineering cares about. Now ask what engineering can do to change it. Maybe the signup flow is slow, or it fails on bad connections, or a confusing step makes people give up. An engineering OKR then targets that: "raise the share of users who complete signup on their first session from 55 to 75 percent." The engineering team has a real way to change a result the whole company wants. That is different from "redesign the onboarding screens," which is a task that may or may not affect the number at all.
This backward direction is what keeps engineering goals honest. It is tempting to start from the engineering side, from the features the team is excited to build, and then pick a metric to justify them. Resist that. Start from the outcome, and let it tell you what to build. Sometimes the answer is a feature. Sometimes it is fixing something slow or broken that no roadmap ever listed. Either one can serve the outcome, and your OKRs should not prefer one. Knowing which of these will actually change the result is its own skill, one we wrote about in knowing what to build.
The mistakes that quietly ruin OKRs
Even teams that understand the outcome idea make a handful of mistakes. They are worth naming so you can spot them.
The first is the to-do list that only looks like a set of goals, which we have already covered: key results that count work instead of results. It is the most common by far.
The second is the guaranteed win. Some teams, worried about missing, set goals they fully control and are certain to hit. "Deploy the new service to production" will happen because the team decides when to deploy. A key result you are sure to hit is not measuring anything. Good OKRs have real uncertainty in them, because a goal you might miss is the only kind that tells you something when you hit it.
The third is tying OKRs directly to pay. The moment hitting a target decides someone's bonus, people stop setting honest targets. They set easy ones they know they can reach, and they hide the ambitious goals that might have taught the company the most. You cannot have both honest goal-setting and OKR-linked bonuses. Pick honesty. Judge people on their real contribution and judgment, which is always broader than one metric.
The fourth is too many objectives. The entire value of OKRs is focus, the act of saying these few things matter most this quarter. A list of ten objectives is not a set of priorities. It is a to-do list with a nicer name, and it tells the team that everything is important, which means nothing is.
Measuring what is hard to measure
A fair objection to all of this: what if you cannot cleanly measure the outcome you care about? Sometimes the thing that matters, like whether users trust the product more, does not come with a simple number.
The answer is to find the closest honest signal and be upfront that it is a stand-in. If you cannot measure trust directly, maybe you can measure how many users return, or how many complete a sensitive action they used to abandon. A rough measure of the real thing is better than a precise measure of the wrong thing every time. And if you truly cannot find any signal for an outcome, treat that as information: it usually means you do not yet understand the outcome well enough to pursue it, and getting to that understanding is valuable work in itself.
Set these each quarter. A quarter is long enough to actually change a real outcome and short enough to change direction when you are wrong. Check in often, every couple of weeks at least, so a goal that is falling behind gets attention while there is still time to do something, instead of showing up as a surprise failure at the end.
Put all of it together and good engineering OKRs stop being a planning ritual and start being a tool that guides the team's work. They direct the team toward a real change in the world, they admit honestly whether the work caused that change, and they keep the number of things you care about small enough that focus is possible. That is harder than listing the features you plan to release. It is also the only version worth doing.
Generated code makes the argument in this post more important, because the metrics it warns about are now easy to increase. A team that is measured on released features or closed tickets can produce an enormous number of both without much difficulty, and the number will look excellent.
Which leaves outcome measures as the only ones that mean anything. Did the process get faster, did the errors go down, did the customer stop complaining. Those were always the better goals. The difference is that the numbers that look good but mean little used to be at least loosely correlated with effort, and that correlation is gone.


