Compliance

Operational Resilience: Important Business Services, Impact Tolerances and Severe but Plausible Testing

Regulatory Counsel · Published August 2026 · Last reviewed August 2026 · 12 min read

Key Takeaways

  • Operational resilience is assessed at the level of the service delivered to the customer or the market, not at the level of individual systems.
  • An impact tolerance is the maximum tolerable disruption to an important business service, expressed in a measurable way, usually time.
  • Mapping must go to sufficient depth to identify the specific people, processes, technology, facilities and third parties on which the service depends.
  • Scenario testing must be severe but plausible, and the value lies in the vulnerabilities it exposes, not in passing the test.
  • Third-party and intragroup concentration risk is a persistent supervisory focus, particularly where several services depend on one provider.
Network of illuminated fibre optic cables in a dark server room representing operational resilience and third party dependencies

Operational resilience reframed a familiar question. Instead of asking whether a firm can recover its systems, it asks whether the firm can continue to deliver the services that matter to customers and to the market, within a tolerance the firm has set and the board has approved.

Step one: identify important business services

An important business service is one whose disruption could cause intolerable harm to customers or risk to market integrity. The test is externally facing. Payroll matters to the firm but is not an important business service. The ability of a customer to access funds, make a payment, place a trade or make a claim generally is.

Common errors at this stage shape everything that follows.

  • Defining services as internal functions or systems rather than customer outcomes.
  • Defining too many services, which dilutes the analysis, or too few, which leaves harm unaddressed.
  • Ignoring services delivered wholly through third parties or intragroup arrangements.

Step two: set impact tolerances

An impact tolerance is the maximum tolerable level of disruption, set at the point at which further disruption would cause intolerable harm. It must be measurable, and time is normally the primary metric, supplemented where relevant by volume, value or number of customers affected.

ComponentWeakDefensible
Metric"Minimal disruption"4 hours from detection
BasisAssumed from IT recovery capabilityDerived from customer harm analysis
ApprovalSet by technology functionBoard approved, with rationale minuted
ReviewStaticReviewed on change of service, provider or volume

The critical discipline is that the tolerance is derived from harm, not from current capability. Setting the tolerance at whatever the firm can currently achieve defeats the purpose, because the gap between tolerance and capability is precisely the information the board needs.

Step three: map the supporting resources

Mapping identifies the people, processes, technology, facilities, information and third parties necessary to deliver each service. Depth matters: a map that stops at "core banking platform" does not reveal that a single upstream data feed, provided by one vendor without contractual resilience commitments, is a point of failure for three services.

Third-party mapping should record the provider, the service supported, contractual resilience terms, substitutability, exit arrangements and any fourth-party dependency the provider itself relies on.

Step four: scenario testing

Scenarios must be severe but plausible. Plausible does not mean comfortable. Useful scenario sets typically include a prolonged outage of a critical third party, a cyber incident causing data integrity loss rather than simple unavailability, loss of a key facility or region, failure of a payment scheme connection, and simultaneous loss of two dependencies.

Testing should record what happened, whether the service remained within tolerance, what the actual recovery path was, and what vulnerability was exposed. A programme in which every test passes is not evidence of resilience. It is evidence that the scenarios are too mild.

Step five: remediation and board reporting

Where testing shows the firm cannot remain within tolerance, that is a vulnerability requiring a remediation plan with an owner, investment where needed and a target date. The board should receive the service list and tolerances, the results of testing, the vulnerabilities identified and the status of remediation, and should apply recorded challenge.

Operational resilience also connects directly to wind-down planning and to prudential assessment: a firm that cannot deliver its services cannot execute an orderly wind-down either.

About Regulatory Counsel

Regulatory Counsel advises UK and international financial services firms on authorisation, prudential and conduct requirements, governance, financial crime and regulator engagement.

Our operational resilience work covers identification of important business services, impact tolerance setting and board rationale, resource and third-party mapping, scenario design and test facilitation, vulnerability remediation planning, third-party risk frameworks and self-assessment documentation.

Contact our regulatory team at info@regulatorycounsel.co.uk.

This article is provided for general information and does not constitute legal or regulatory advice. Firms should confirm the current position against FCA publications and take advice on their specific circumstances.

Frequently Asked Questions

A service delivered to an external customer or to the market whose disruption could cause intolerable harm to customers or risk to market integrity. It is defined by the outcome delivered externally, not by internal systems or functions.

The maximum tolerable level of disruption to an important business service, expressed in measurable terms, usually a period of time, set at the point beyond which harm becomes intolerable.

They should be derived from analysis of customer and market harm, approved by the board with a recorded rationale, and reviewed when the service, provider, volumes or dependencies change, rather than being set to match current recovery capability.

Testing against disruptions that are demanding yet realistic, such as prolonged failure of a critical third party, a data integrity cyber incident or loss of a payment scheme connection, with the purpose of exposing vulnerabilities rather than passing.

Because multiple important business services often depend on the same provider or the same underlying infrastructure, so a single failure can breach several impact tolerances simultaneously and the firm may have no substitutable alternative.

Need Expert Advice?

Free initial consultation. No obligation.

Speak to an Expert