Operational resilience

Operational resilience, measured in tolerances.

Operational resilience is an organisation’s ability to keep delivering the services its customers depend on through disruption — including disruption it did not cause and cannot prevent. What separates it from business continuity is the starting point: it begins from the service the customer receives and the harm caused when that service stops, not from the systems you happen to own.

Last reviewed by Darren Craig

What is operational resilience?

A change of unit, and everything follows from it.

Operational resilience is the ability to prevent, adapt to, respond to, recover from and learn from operational disruption. The definition is unremarkable. What makes it a distinct discipline is the unit it measures: the service delivered to a customer or the market, rather than the application, the site or the supplier.

That shift has a consequence supervisors were explicit about. A firm can hold individually acceptable recovery times for every system it owns and still be unable to say how long a customer would be unable to make a payment — because the answer depends on systems, people, premises, data and third parties in combination, and nobody owns the combination. Resilience is assessed end to end or it is not assessed.

The second consequence is the one that pulls third-party risk into it. Most of a modern service is delivered by parties the firm does not control. You are accountable for the resilience of a service that runs substantially on somebody else’s infrastructure, and no supervisor has ever accepted that as a reason for not knowing.

The operational resilience framework framework

Four steps, in this order. UK firms have worked to this sequence since the FCA and PRA rules took effect, and it is a sound structure whether or not you are regulated.

  1. 1. Identify important business services
    A service delivered to an external customer or the market, whose disruption would cause intolerable harm to customers or risk to market integrity. Named from the outside in — "making a payment", not "the payments platform". Most firms find they have fewer than they expected, and that the list argues itself into shape faster once harm rather than revenue is the test.
  2. 2. Set an impact tolerance for each
    The maximum tolerable disruption — usually expressed in time — before the harm becomes unacceptable. Not a target, and not what you can currently achieve: a limit set from the customer’s position. An impact tolerance that matches your existing recovery time has almost certainly been written backwards from the capability rather than forwards from the harm.
  3. 3. Map what the service depends on
    People, processes, technology, facilities, information — and third parties, to the depth that matters. Mapping is where the uncomfortable findings live: the single-supplier dependency nobody had registered, the fourth party three of your services share, the manual step one person performs.
  4. 4. Test against severe but plausible scenarios
    Not the disaster-recovery scenario you can pass. Supervisors ask for severe but plausible — a critical supplier offline for days, a ransomware event, the loss of a data centre — run to establish whether you would stay inside the tolerance. A test that always succeeds is measuring the wrong thing.

Operational resilience vs business continuity

They overlap, and the difference is not academic — it changes what gets planned, what gets tested and what "good" looks like at the end.

Business continuity
Recover the organisation
  • Starts from your systems, sites and processes
  • Asks how quickly you can restore what broke
  • Measured in RTO and RPO per system
  • Plans for identified scenarios
  • Succeeds when internal operations resume
Operational resilience
Keep the service running
  • Starts from the service a customer receives
  • Asks how long a customer can be without it
  • Measured in impact tolerance per service
  • Assumes disruption will happen, including from causes you cannot foresee
  • Succeeds when the customer stays inside the tolerance — however that is achieved

Where third parties decide the answer

The dependency you did not build, and cannot patch.

Mapping usually establishes that an important business service depends on a handful of suppliers whose failure you would absorb rather than fix. That turns four questions from third-party risk into resilience questions:

Substitutability. Could this supplier be replaced inside the impact tolerance? For most cloud, payments and identity providers the honest answer is no, and that is worth stating plainly rather than assuming an untested exit plan covers it. Concentration. How many of your services stop if one provider does, including through suppliers that look unrelated on the register?

Visibility. Would you learn about the supplier’s degradation from the supplier, or from your customers? This is the practical argument for continuous monitoring of supplier posture rather than annual assurance. Contractual reach. Do the terms give you notification windows, audit rights and exit provisions that a severe scenario would actually make use of?

Where a firm is regulated, this is also where the vocabulary converges: material outsourcing under the UK supervisory statements, ICT services supporting critical or important functions under DORA, and the critical third parties regime for the providers a whole sector shares. Mapping the classifications once beats maintaining three registers of the same suppliers.

Who requires it

Three regimes, converging on the same four steps.

UK financial services. The FCA (SYSC 15A) and the PRA require firms to identify important business services, set impact tolerances, map dependencies and test against severe but plausible scenarios — and, since the transitional period closed on 31 March 2025, to be able to remain within those tolerances.

The EU. DORA has applied since 17 January 2025, covering ICT risk management, incident reporting, resilience testing, ICT third-party risk and information sharing. Its third-party pillar is the most demanding part for most firms, which is why it has a page of its own.

Elsewhere. APRA CPS 230 takes the same shape in Australia, and NIS2 pushes comparable expectations across essential and important entities in the EU beyond financial services. The regimes differ in scope and vocabulary far more than in method.

Operational resilience software

What tooling holds, and what it cannot decide.

Resilience tooling generally does three things: it holds the service-to-dependency map so it can be queried rather than redrawn, it records scenario tests and their outcomes against tolerances, and it produces the self-assessment a board or a supervisor reads. Firms that run this in spreadsheets do not usually fail because the spreadsheet is wrong — they fail because it is eleven months old.

What no tool decides: which services are important, what harm is intolerable, and what you do about a critical supplier you cannot replace inside the tolerance. Those are judgements, and they are the ones a supervisor asks about. Where RiskXchange contributes is the supplier half of the map — continuous outside-in monitoring of the providers your important business services depend on, so the dependency picture reflects this week rather than the last review.

Operational resilience, answered.

What is operational resilience?
The ability to keep delivering the services customers depend on through disruption, and to recover and learn from it. It is measured per service rather than per system, and it assumes disruption will happen — including disruption caused by third parties you do not control.
What is the difference between operational resilience and business continuity?
Business continuity starts from your systems and asks how quickly you can restore them. Operational resilience starts from the service a customer receives and asks how long they can be without it. One is measured in recovery times per system, the other in impact tolerances per service — and a firm can be strong at the first while unable to answer the second.
What is an impact tolerance?
The maximum tolerable disruption to an important business service before the harm to customers or the market becomes unacceptable — normally expressed as a duration. It is set from the customer’s position, not from what the firm can currently achieve. A tolerance that matches your existing recovery capability was probably written backwards.
What is an important business service?
A service delivered to an external customer or the market whose disruption would cause intolerable harm or threaten market integrity. Named from the customer’s side — "making a payment" rather than "the payments platform" — which is what stops the list becoming an inventory of systems.
Who regulates operational resilience in the UK?
The FCA and the PRA, through SYSC 15A and the PRA’s operational resilience policy. Firms must identify important business services, set impact tolerances, map dependencies and test against severe but plausible scenarios; the transitional period for remaining within tolerances closed on 31 March 2025.
How does operational resilience relate to third-party risk?
Most of an important business service is usually delivered by third parties, so mapping ends at suppliers you cannot fix. The resilience questions become substitutability inside the tolerance, concentration across providers, whether you would notice degradation before your customers do, and whether the contract gives you anything useful in a severe scenario.

Your map is only as current as your suppliers are.

See continuous monitoring on the providers behind one of your important business services.