When to rewrite vs refactor legacy code (and why the big rewrite is usually wrong)

At some point every team inherits code that is hard to work with. It is slow to change, easy to break, and nobody fully understands it anymore. The people who wrote it have moved on, and the ones who are left approach it carefully, because any change might break it. Sooner or later someone says the obvious thing: we should just rewrite this. Start fresh, do it right, be finished with the old code. It sounds clean and honest. It is almost always the wrong instinct, and the reason is not obvious until you have watched a rewrite fail.
Here is how we think about the choice between rewriting old code and improving it in place, and how to replace a legacy system without stopping the business to do it.
Why the big rewrite is almost always wrong
The case for a rewrite feels certain. The old code is ugly and slow to change. A fresh version would be clean, fast, and built the way you would build it today. Why keep improving something badly broken when you could start over?
Because the ugly old code is not just ugly. It is full of years of small fixes for real problems, most of which are invisible until you delete them. That strange condition nobody understands is there because a specific customer hit a specific bug in 2022. That odd extra step in the checkout flow is handling a tax rule in a country you forgot you sell in. None of it is documented, because it accumulated one urgent fix at a time. A rewrite starts from nothing and has to rediscover every one of those unusual cases through costly mistakes, usually by breaking them in front of customers and getting the same bug reports the old team already fixed.
Meanwhile, the business does not stop needing new things. So the team building the rewrite is trying to match a system that keeps changing: they have to match everything the old system already does, which took years to build, while the old system keeps adding more. Most rewrites run late for exactly this reason, and many get abandoned before they ever replace the original, after months or years of work. And the whole time, the old system usually gets no major features, because your best people are busy replacing it. If the rewrite runs a year late, you have lost a year of forward progress. This is the same danger we described in build vs buy: when custom software is actually worth it: the bigger the build, the more places it can go wrong, and a from-scratch rewrite is one of the biggest builds a team can take on.
Why engineers keep choosing it anyway
It helps to be honest about where the urge comes from. Reading and understanding someone else's code is hard, slow, and not much fun. Writing fresh code is faster and far more satisfying. So when an engineer looks at a legacy system, "let's rewrite it" is partly a technical judgment and partly a wish to avoid the unpleasant work of learning why the old code does what it does.
That preference is human, and it is not shameful. But it is not a business reason, and the trouble starts when the two get confused. "This would be more pleasant to rewrite" quietly becomes "this needs to be rewritten," and a preference gets presented as a necessity. Before you commit a team to a rewrite, it is worth asking plainly: are we doing this because the foundation genuinely has no future, or because reading the old code is annoying? The answer changes the decision.
When a rewrite is genuinely justified
Rewrites are not always wrong. Sometimes the old system really has no future, and no amount of careful cleanup fixes that.
The clearest case is a foundation with no future. The code runs on a platform that is no longer supported and will not get security updates. It is written in a language or framework nobody will build a career on, so you cannot hire people to maintain it. It depends on a database or a service that is being shut down. When the platform under the code is going away, improving the code does not help. You need a new platform.
The other case is when the cost of safe change has become higher than the cost of replacement. If every small change to the old system risks breaking three things you did not touch, and every fix takes weeks of careful work because the code is difficult to change at every step, then at some point replacement is genuinely cheaper than continuing. But be strict with yourself here. This is a real threshold, not a feeling, and most code that annoys people has not actually reached it. Even when a rewrite is justified, "rewrite" should almost never mean "shut everything down and rebuild it all at once." It should mean replace it piece by piece, which is the next part.
Refactor first, in small safe steps
Before you choose replacement, remember the other option that gets skipped in the excitement of a fresh start. You can improve the existing code in place, in small steps, while it keeps running in production. That is refactoring, and for most legacy problems it is the right answer.
The discipline is doing it small. You change one piece, test it, release it, confirm nothing broke, then move to the next. The system keeps running and stays useful the entire time, so you are never in the dangerous position of having a half-built replacement that does not fully work. This is the same reason strong teams release in small increments generally: the long-running DORA research keeps finding that small, frequent changes go with lower failure rates than big, rare ones. Refactoring in small steps applies that same lesson to old code. You are not risking the whole system on one big change. You are cleaning up the old code in small amounts, and you can stop whenever the code is good enough, which is a luxury a rewrite never gives you.
Replace piece by piece with the strangler fig approach
When the foundation really is dead and replacement is genuinely needed, there is still a safe way to do it that is not an all-at-once rewrite. It is called the strangler fig approach, named after a vine that grows around a tree until the tree is gone and the vine stands on its own.
The idea is simple. You build the new system beside the old one, not instead of it. Then you move one small piece of work over: route a single feature, or a share of traffic, to the new code. You confirm it works in production with real users. Then you move the next piece, and the next, while the old system keeps handling everything you have not migrated yet. Over weeks and months, the new code surrounds the old code until the old system has nothing left to do and can be switched off.
This changes everything about the risk. The business never stops, because the old system keeps running the whole time and you keep releasing new features in the new code as you go. Nothing depends on one giant switch-over that has to work perfectly on a Friday night. If a migrated piece has a problem, it is one small piece, easy to see and easy to roll back, not a hundred changes hidden inside one release. And you can pause or reprioritize at any point with real value already delivered, instead of spending a year on a rewrite that is worth nothing until the day it finally replaces the original. When we help a team move off a system with no future, this is almost always the plan, because it lets us move fast without breaking the product.
How to decide honestly
Put it together and the decision comes down to one clear question: is the problem the structure of the code, or the platform it runs on? If good engineers could improve the code in small steps while it keeps running, refactor. That covers most legacy problems, even the code that feels impossible to fix. If the foundation genuinely has no future, a platform that will not be supported, a dependency being shut off, a cost of change that has truly become higher than the cost of replacement, then plan a replacement, and plan it as many small piece-by-piece migrations, not one large rewrite.
And before you commit either way, be honest about the motive. A rewrite because the platform is dead is a business decision. A rewrite because reading the old code is unpleasant is a preference presented as a business decision. The teams that stay out of trouble are the ones that can tell the difference, and that choose the boring, gradual path far more often than starting from nothing. Starting from nothing feels like the responsible choice. Most of the time, the responsible choice is to keep the old system running and quietly replace it one part at a time.
What cheap regeneration does to this decision
The rewrite argument deserves revisiting, because the thing that always made rewrites fail was cost: they took a year, the business kept moving, and the new system never caught up. AI code generation shortens that year, which makes a full rewrite genuinely viable in cases where it never was.
It does not make it safe. The reason a rewrite is risky was only partly the time it took. The rest is that nobody fully knows what the old system does, and the undocumented behaviour is important in ways that appear at the switch-over. That problem is untouched by faster building. If anything it gets worse, because a rewrite you can produce in three weeks is a rewrite you might attempt without doing the characterisation work first.
So the rule changes but does not reverse. Before, the answer was almost always refactor, because rewriting cost too much. Now: characterise the old system first, build a behavioural test suite from what you find, and treat that suite as the specification. If you have done that, a rewrite is a reasonable option in a way it was not before. If you have not, faster tooling has only made it cheaper to build the wrong thing.
Sources
- DORA, Accelerate State of DevOps research: https://dora.dev/research/
- McKinsey and University of Oxford (2012), "Delivering large-scale IT projects on time, on budget, and on value": https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/delivering-large-scale-it-projects-on-time-on-budget-and-on-value


