Operational Resilience Lessons from the NATS Outage.

This week, UK air traffic control hit serious trouble, again, after a fault in NATS's flight-processing system led to approximately 2,000 cancellations and stranded tens of thousands of passenger. Heathrow bore the brunt, with Gatwick, Stansted, Birmingham and Manchester all significantly affected too, though NATS was clear that UK airspace itself stayed open throughout. Recovery ran into a second day, and both Ryanair and WizzAir chiefs have echoed the public backlash. This is the third major NATS technical disruption in just over three years.

NATS operates one of the most safety-critical systems on the planet, and the people on the ground worked hard to keep everyone safe while things unravelled. 

The specific trigger behind this week's disruption remains under investigation. What we do know is that the fault affected NATS's flight-processing systems, forcing restrictions on traffic and creating widespread disruption across UK airports. For business leaders, though, the more interesting question isn't the precise technical cause. It's what happens when a critical system fails, and whether the resilience measures around it perform as expected.

That's where the 2023 incident remains relevant. The independent investigation found that when the secondary system automatically took over, it encountered the same defect and shut down too. Both systems ran the same software, meaning the same flaw existed in both environments. The result was a common-mode failure: redundancy was in place, but both systems shared the same vulnerability.

Most people won't remember this incident as "NATS had a software fault." They'll remember it as cancelled flights, stranded families, a CEO facing public calls for his resignation within a day of the disruption, and a bill that airlines are still totting up. The technology failure was the trigger. The consequences were entirely business ones: customer experience, reputation, financial exposure, and now regulatory and government attention on whether the organisation is fit for purpose.

 

Operational resilience has stopped being a technology conversation and become a governance one. It's no longer just "does IT have a backup plan"; it's "can this organisation demonstrate, credibly, that it would keep functioning under stress." One is a project. The other is something you have to be able to prove, on demand, to a board, a regulator, or a customer who's asking harder questions than they used to.

 

The real lesson isn't "you need a backup." It's "you need proof it works."

NATS had a backup system. That's exactly the point. Nobody sat in a planning meeting and decided not to bother with resilience: they built it, funded it, signed it off, and by NATS's own account put it through comprehensive testing. The independent panel that later investigated found something more specific and, honestly, more relatable: a piece of software logic needed to handle exactly this kind of duplicate-code scenario had been missed during development, and wasn't caught by assurance or acceptance testing before go-live. Every possible combination of routing and data conditions genuinely can't be tested for, but this particular one slipped through, in both the primary system and the backup that was supposed to catch what the primary missed.

That's the uncomfortable, useful lesson: a resilience plan that exists on paper and one that's been proven against realistic, awkward, edge-case scenarios are different things, and only one of them tells you anything true about what will actually happen on a bad day. "We have DR" is a sentence. "We've tested it against the scenarios that actually break things" is a capability. Increasingly, that's the distinction customers, boards and regulators are actually asking about: not whether resilience exists on paper, but whether anyone can show it holds up under the kind of unusual, unglamorous edge case that's easy to miss and, by definition, is exactly the one that eventually happens.

 

It's not just your own infrastructure anymore, either

There's a second layer to this that's easy to miss if you're picturing resilience as something that happens inside your own four walls. Most organisations today don't run on their own infrastructure alone: they run on Microsoft, on cloud platforms, on connectivity providers, on security vendors, on a chain of suppliers they didn't build and don't fully control. A failure at any one of those can become your outage, your customer service failure, your reputational problem, with your name on the apology email. The question worth asking isn't just "is our infrastructure resilient"; it's "is our whole ecosystem resilient," including the parts you've outsourced and the parts you've simply never stress-tested.

And it's worth saying plainly: this lesson doesn't only apply to accidental technical failures like NATS's; NATS itself has said a cyberattack has been ruled out here. But the same gap, a resilience plan that's never been properly tested against a real scenario, is exactly what turns a ransomware incident, a compromised account, or a supplier breach into a prolonged business shutdown rather than a contained one. Cyber resilience and operational resilience are the same conversation now, not two separate ones.

 

Underneath all of it: technical debt, and a bill that doesn't go away

This wasn't really a "one duplicate waypoint" problem, not entirely. NATS operates within an air traffic infrastructure that's been publicly criticised for its age before: commentary around a separate 2014 outage noted that the underlying flight-data processing system has software heritage going back to 1960s-era US systems, with newer interfaces layered on top ever since.

To be fair, that's context rather than a smoking gun: nobody has established that 1960s-era code caused either the 2023 or 2026 faults. But it's a useful reminder that in infrastructure this long-lived and safety-critical, old foundations and new logic have to coexist for decades, and that coexistence is exactly where technical debt tends to hide.

NATS's own history shows that debt compounding in real time. That 2014 outage was traced to a latent fault in a data-type definition, buried in roughly two million lines of code in a different system dating from the 1990s, not, as it's sometimes simplified, "a single line of code," though the effect was similarly disproportionate: one flawed assumption, a day of national disruption, and a government minister publicly calling the infrastructure ancient. A five-year, multi-hundred-million-pound modernisation programme followed. Delivery on replacing the ageing systems has since been repeatedly pushed back: NATS's own reporting points to a mix of legacy-system complexity, technical replanning and operational transition constraints, alongside budget prioritisation, as reasons the full replacement keeps slipping. Then 2023 happened. A radar-related fault in July 2025 followed, resolved within around 40 minutes via backup, with roughly 150 cancellations, considerably less severe than 2023 or this week. Then this week, with the flight-processing fault's root cause still under investigation. That's not simply a story about bad luck, and it's not purely about money either: it's a story about technical debt, resilience investment and governance decisions that keep landing in the "important, not urgent" pile, one budget cycle at a time, until urgent turns up on its own schedule. The bill doesn't disappear. It gets bigger, and it turns up at a worse time, usually as a board-level crisis rather than a line item.

Most mid-market and enterprise leaders we talk to recognise this pattern immediately, because it isn't unique to national infrastructure. It's the legacy system nobody wants to touch because the person who built it left years ago. It's the supplier dependency nobody's mapped. It's the "we'll test failover properly once things calm down" that's been true for two years. It's rarely one big dramatic decision to run outdated, unproven infrastructure: it's a hundred small, reasonable-sounding decisions not to fix it, map it, or test it this quarter, that quietly stack into exactly the kind of fragile foundation that turns a minor edge case into a very bad Tuesday, or a very bad board meeting.

 

What operational resilience actually needs to mean

Not a tickbox. Not a line in a contract nobody's read since it was signed, and not just an IT project. Real operational resilience means being able to answer, with evidence:

  • Have we actually tested our recovery and failover capability against realistic failure scenarios, not just documented that one exists
  • Do we know where our real dependencies sit, across suppliers, cloud platforms and vendors, not just inside our own infrastructure
  • Have we stress-tested the extreme, unlikely scenarios, not just the routine ones, because those are the ones that eventually happen
  • Is cyber resilience part of the same plan as operational resilience, rather than a separate conversation with a separate owner
  • Could we show this to a board, a regulator or a customer today and prove it, not just describe it

We've been doing this since 1985. We've watched technology change more times than we can count, and the one thing that hasn't changed is this: the organisations still standing, and still trusted, when things go wrong are the ones that treated resilience as a capability to prove, not a project to file away. Because things will go wrong. That's not pessimism, it's just how business works now. The difference is what happens in the next five minutes, and whether anyone can actually stand behind what was promised on paper.

Running a business is hard enough without wondering what your version of "the identically named waypoint" is, or which supplier, system or untested assumption is one bad day away from being it. That's exactly the conversation we have with our clients: helping them tackle technical debt, map supplier and ecosystem risk, and build operational, cyber and disaster recovery resilience that can actually be demonstrated, not just described.

If you couldn't prove your recovery capability to a board tomorrow, that's worth a conversation before something forces the issue.

[Talk to us about operational resilience →]