Modernising a system you cannot switch off
Big-bang rewrites fail for predictable reasons. A safer route replaces a legacy system one capability at a time, while it stays in service.
7 min read
Most organisations reach the same conclusion about their oldest system: it needs replacing. The instinct that follows is usually to rebuild it completely, run the two systems side by side for a short period, and switch over on an agreed date. That plan is attractive because it is easy to describe. It fails often enough that it deserves more scepticism than it usually gets.
Why the rewrite keeps failing
A working system that has run a business for a decade contains more knowledge than anyone remembers putting there. Its behaviour encodes exceptions, informal rules, and workarounds that were never documented because they emerged one support call at a time. A rewrite starts from the specification people can articulate, which is a much smaller thing than the system that actually exists.
- Scope is estimated against the described system, not the real one, so it grows continuously
- The old system keeps changing during the rebuild, because the business has not stopped
- Nothing is delivered until the end, so there is no early feedback and no way to course-correct cheaply
- The switch-over concentrates every risk into a single day
The last point matters most. If the migration goes badly, the fallback is to revert to the legacy system, which means the entire investment sits unused while confidence in the project drains away.
Replace one capability at a time
The alternative is to put a routing layer in front of the legacy system and move functionality across in slices. Each slice is chosen so that it can be delivered, tested, and switched independently. The old system keeps serving everything not yet moved, and shrinks as the new one grows.
- 01Put an interface in front of the existing system so callers no longer depend on its internals
- 02Pick the first slice on the basis of risk and value, not on which part is most interesting to build
- 03Build and run the new implementation alongside the old one, comparing outputs before switching traffic
- 04Move traffic gradually, keeping the ability to route back
- 05Retire the old code path only once the new one has run in production long enough to be trusted
The goal of each step is not to finish the migration. It is to leave the system in a better state than you found it, with the option to stop.
What this costs you
This approach is not free. Running two implementations of the same capability, even briefly, means maintaining both and building the comparison tooling to check that they agree. Teams sometimes read that overhead as waste. It is better understood as the price of being able to stop safely at any point, which is precisely what a big-bang cutover does not offer.
It also asks something harder of the people commissioning the work: accepting that the system will spend a long period in a hybrid state. That is uncomfortable to explain to a board expecting a launch date. The compensation is that value arrives throughout, rather than all at the end or not at all.
When a rewrite is genuinely the right call
Incremental replacement is not always correct. If the existing system is small enough to be rebuilt in a few weeks, the overhead of running both is not worth it. If the underlying platform is genuinely unsupportable, or the data model is wrong in a way that every new feature has to work around, a clean start may be cheaper. The distinction is size and structural soundness, not age. Old and boring is not the same as broken.