Most legacy systems need less than their owners fear and more than they budget. A full rewrite is rarely the right answer, and leaving a system alone stopped being free once its runtime and framework started to lose security support.

For a single system the decision comes down to four questions, asked in order. The stories below, from airlines, banks, governments and our own clients, show what a wrong answer cost, and what a good one looked like.

A system becomes legacy when changing it costs more than the change is worth

Age alone does not make a system legacy. A ten-year-old service with good tests, current libraries and three engineers who know it well is an asset. A three-year-old one that only its author can change, on a framework that has left support, is already legacy.

The US federal government is the largest published example of what that costs over time. In a 2016 review, the Government Accountability Office found that about 75% of the more than $80 billion federal IT budget for fiscal year 2015 went to operating and maintaining existing systems, and that 5,233 of roughly 7,000 IT investments spent all of their money on operations and maintenance, with nothing left for improvement (GAO-16-468). One of the systems it named was the IRS Individual Master File, written in assembly language and holding data from about a billion taxpayer accounts. It was still processing refunds when the report came out.

A very old system can be the right one to keep, as long as the business rarely needs it to change and someone can still keep it safe. The trouble starts when either of those stops being true.

Keeping a system has a price, and it arrives all at once

In the last week of December 2022, winter storms paralysed Southwest Airlines' operations in Denver and Chicago, and the airline's crew rescheduling could not keep up. Southwest cancelled nearly 17,000 flights and stranded more than two million travellers over the holidays. The company later put the cost at more than $1.1 billion in refunds, reimbursements, extra costs and lost ticket sales, and in December 2023 it agreed a $140 million settlement with the US Department of Transportation (CBS News). The crew scheduling tool, SkySolver, was an off-the-shelf application customized for Southwest and in use for decades (Texas Standard). At the Senate hearing in February 2023, the president of the pilots' union testified that pilots had warned about the scheduling technology and processes for years (written testimony). That is the union's account; Southwest's chief operating officer also told the committee that the scheduling system had been overwhelmed (Bloomberg).

In April 2020, New Jersey saw unemployment claims rise by 1,600%, with more than 206,000 new claims in a single week. The benefits system ran on COBOL, and Governor Phil Murphy said in a press briefing that the state needed people with COBOL skills alongside health workers (StateScoop). The system had worked for decades at normal volumes. What failed first was the number of people who could change it quickly.

Both systems ran well for years, and in both cases the cost came due under peak load, in one week, with no time to plan. The risk in keeping a legacy system comes from three things together: code that is hard to change, few people who can change it, and a day when it has to change fast.

Four options, and four questions that pick between them

There are four things you can do with a legacy system, and they differ in cost and risk by an order of magnitude.

OptionTypical costTime to valueMain riskWhat you keep
Keep and containLow: patching, monitoring, documentationImmediateThe skills and support gap grows each yearEverything
Upgrade in placeWeeks per applicationWeeksLow with tests, high without themThe code, the data and the behaviour
Rearchitect in slicesMonths, spread over releasesAfter the first sliceA long middle with two systems side by sideThe old system, until each slice replaces part of it
ReplaceA year or moreAt cut-overOne large cut-over, and scope that grows as hidden rules surfaceThe data, if it migrates cleanly

The four questions below pick between them. The first asks whether the business still needs the system to change, because that decides whether its design matters at all.

A decision tree for one legacy system. Question 1: will the business need to change the system in the next two years? If no, question 2: is the system on unsupported or unpatched versions? If yes, upgrade it in place; if no, keep and contain it. If the business will need changes, question 3: does the system’s design block those changes? If no, upgrade it in place. If yes, question 4: could a packaged product, bought rather than built, replace the system? If yes, replace it in stages; if no, rearchitect it in slices. Every yes leads to a larger change.QUESTION 1Will the business need to changethe system in the next two years?QUESTION 2Is the system on unsupportedor unpatched versions?QUESTION 3Does the system’s designblock those changes?Keep and containPatch and monitorUpgrade in placeSame designQUESTION 4Could a packaged productreplace the system?Rearchitect in slicesOne part at a timeReplace in stagesSegment by segmentNoYesNoYesNoYesNoYes
Follow the questions from the top for each system. Most systems end at upgrade in place, by one route or the other.

Every question is phrased so that yes means a larger change. The first splits systems by whether the business will ask for changes; the change requests of the last twelve months are the best evidence, and a roadmap the second best. A system nobody needs to change only has to be safe, which is the second question: if its runtime or libraries are out of support or unpatched, upgrade it in place without touching the design, and if not, keep and contain it. For a system the business will keep changing, the third question asks whether its design blocks those changes. Most designs do not, and then an upgrade in place, repeated as new versions arrive, is enough. When the design does block the work, such as a batch process that has to become real time or a data model that no longer fits how the business works, the last question is whether a packaged product could now replace the system: software you buy or subscribe to, such as a payroll, CRM or policy administration suite, used in its place. If one could, move to it in stages. If not, rebuild it one slice at a time with the strangler fig pattern.

Keep and contain still takes work. Containing a system means putting it behind a stable interface so that new code talks to the interface and not to its internals, keeping the runtime and the operating system patched, monitoring the few things that would hurt if they broke, and writing down what it does while the people who know are still around. It also means a date in the calendar, once a year, to ask the four questions again.

This tree is for one system. When a company has twenty, the order in which to work through them matters as much as the answer for each, and our guide to application modernization covers how to sequence a portfolio against the engineers you have.

Rewrites and big-bang migrations fail big

Bent Flyvbjerg and Alexander Budzier studied 1,471 IT projects and found an average cost overrun of 27%, with one project in six overrunning by 200% on average (Harvard Business Review). The two stories below come from that tail.

In April 2018, TSB moved the data of its customers and its corporate services onto a new IT platform. The data migrated, but the new platform failed as soon as it went live. All of TSB's branches and a large share of its 5.2 million customers were affected, some for months, and the bank did not return to business as usual until December 2018. In December 2022 the Financial Conduct Authority and the Prudential Regulation Authority fined TSB £48.65 million, finding that it had failed to organise and control the migration programme adequately and to manage the risks of its outsourcing to its critical supplier. By then TSB had paid £32.7 million in compensation to customers (FCA).

Queensland Health replaced its end-of-life payroll system, LATTICE, in 2010. The prime contract, awarded to IBM in December 2007, was planned at about $98 million. The project cost $181 million by its end, and the estimated cost of fixing, maintaining and running the system came to about $1.2 billion over eight years. The commission of inquiry concluded that the replacement "must take a place in the front rank of failures in public administration in this country" (Queensland Health Payroll System Commission of Inquiry, 2013).

Both were replacements of systems that were at the end of their life, so doing nothing was not an option. In both, the change came as one cut-over that carried all the risk at once, with no way back once it had started.

When the answer is replace, replace in stages

If the tree ends at replace, the TSB and Queensland stories still apply, and most of their risk can be designed out. The practices below come from our own migrations and from the post-mortems above:

  • Move by segment, not all at once. A product line, a region or a group of customers goes first, and the rest follow once it is stable.
  • Run old and new side by side and compare the results. For a payroll, that means parallel pay runs and a diff of the payslips; for a trading or billing system, the same inputs through both and a reconciliation of the outputs.
  • Rehearse the data migration until it is dull. Run it end to end on a copy of production data several times, time it, and fix every record it rejects before the real weekend.
  • Decide how to go back before you go forward. Write the rollback plan, test it, and agree in advance what would trigger it.
  • Keep the old system readable after cut-over. Reconciliation questions arrive for months, and the answers are in the old data.

Old code that nobody uses is a risk of its own

Keeping a system does not mean keeping everything in it. On 1 August 2012, Knight Capital deployed new trading code to its order router. The new code reused a flag that years earlier had switched on a function called Power Peg, which Knight had stopped using in 2003 but never removed. A technician copied the new code to seven of the eight servers. On the eighth, the old flag woke the old code, which sent millions of orders into the market. Knight lost more than $460 million in about 45 minutes, and before the market opened its systems had sent 97 automated emails about the error that nobody acted on (SEC order, 2013).

Unused code, dormant feature flags and half-retired integrations all look harmless because nothing calls them. They still ship with every deployment. An inventory of them, and a plan to delete them, belongs in every modernization, including the ones that keep the system.

When a rewrite is right, run it next to the old system

Some systems do need to be rebuilt, and the difference between the TSB story and a good rebuild is mostly in how the switch is made.

In 2019 Shopify set out to rewrite the part of its platform that renders merchants' storefronts, which had grown inside a Ruby on Rails monolith over fifteen years. The team kept the old implementation serving customers and built the new one next to it. A verifier service sent samples of real production traffic to both and compared the responses. Once an endpoint for a given shop had matched often enough, routing started sending that shop's traffic to the new renderer, and any request the new code failed to render fell back to the old one. The team started with simple pages and moved to product and collection pages later. Average server response times on the new implementation came out four to six times faster than the legacy one (Shopify Engineering).

A rebuild does not have to replace the whole system either. A UK hotel bookings aggregator came to us when its availability search had outgrown a cluster of dozens of SQL Server machines: adding one more raised cost by about 2% and throughput by about 0.1%. We kept SQL Server as the system of record and moved availability search into an in-memory data grid, loaded from the database and kept current by a message queue. Traffic grew from about 200 million to more than 300 million queries a day, response times fell by 60%, and infrastructure cost fell by 80% (case study). The rest of the system stayed as it was. Our Performance Rescue work usually starts from this kind of problem.

Capacity, not technology, is the constraint for most teams

The people who understand a legacy system are almost always the same people building features on it, and in many companies there are one or two of them. That limits every option: an upgrade, a rearchitecture and a replacement all need their time, and the business rarely agrees to stop feature work while they do it.

The options that fit around feature work are the ones that move in small steps. Earlier this year we upgraded a fifteen-year-old online store and back-office system for a wine merchant from Java 17 and Spring 5.3 to Java 25 and Spring 6.2, in about four weeks and without a feature freeze, and cut its critical and high-severity advisories from 47 to none (case study). It was an upgrade in place, not a rewrite. A system that only one engineer can change is a risk even when every version is current, and pairing a second engineer on the next upgrade is often the cheapest fix on the list.

What the first two weeks should produce

Before choosing an option, get a clear picture of what actually runs in production. For a single system, two weeks is usually enough to produce:

  • An inventory of what runs: runtime and framework versions, scheduled jobs, integrations, and every system that calls it.
  • The support status of each component against its vendor's dates, and a dependency scan with the open vulnerabilities.
  • The change requests of the last twelve months, which answers the first question with data instead of opinion.
  • The five or so flows the business depends on, and whether each has an automated test.
  • A list of dead code, dormant feature flags and unused integrations, with a plan to remove them.
  • The answers to the four questions, the option they point to, and a written estimate for it.

In an acquisition the same questions come up with a deadline attached, and they are the core of our technical due diligence.

If you would like a second opinion on one of your own systems, that is what the first step of our Application Modernization Sprint produces: a written view of which option the system needs, the order of work and a ballpark estimate, within a week of a thirty-minute call.