Operational resilience is not only the recovery of servers or buildings. Banking leaders need to know which customer and business outcomes are most important, how long or how severely they can be disrupted, and whether the combined people, process, technology, data, facility and third-party arrangements can keep those services within the bank’s tolerance.

01

Leaders define the service from its outcome

A critical business service is described as an outcome the bank must deliver, such as customers accessing funds or the bank completing an essential payment obligation. Naming a system, department or vendor alone is too narrow because the service usually crosses several organizational and technology boundaries.

Leaders identify services whose disruption could materially harm customers, the bank or the broader financial system, consistent with the institution’s size, activities and risk profile. A disciplined inventory prevents every activity from being labeled critical while also preventing a familiar internal hierarchy from obscuring an essential customer-facing dependency.

02

A tolerance turns importance into a decision boundary

For each service, governance defines the disruption the bank is prepared to withstand, considering duration, volume, data integrity, customer impact and other relevant measures. Recovery-time targets can support this view, but a technology target alone may not describe the point at which the complete service causes unacceptable consequences.

The tolerance is approved and reviewed at the appropriate level and aligned with risk appetite, legal and contractual obligations and credible capabilities. It guides design and investment; it is not permission to ignore harm until a timer expires.

03

Mapping exposes the dependencies behind the outcome

Teams map the people, procedures, applications, infrastructure, data, facilities and third parties required to deliver the service from beginning to end. They also identify handoffs, shared resources and upstream or downstream services whose failure could interrupt the same outcome.

The map should support decisions rather than become a static diagram. Owners keep it connected to inventories, change management, vendor oversight and incident records so a system replacement, staffing change or new concentration updates the resilience view before the next test or disruption.

04

Severe but plausible scenarios test the whole service

Exercises combine hazards such as cyberattack, data corruption, facility loss, telecommunications failure, third-party outage or unavailable personnel. The objective is to test whether the service can remain within tolerance when several supporting assumptions fail, not merely whether one component can restart in isolation.

Tests examine detection, decision rights, workarounds, capacity, customer communication, reconciliation and recovery of accurate data. Independent challenge and evidence from actual incidents help leaders distinguish a demonstrated capability from a plan that works only on paper.

05

Governance converts weaknesses into prioritized action

Results show where a single point of failure, concentrated provider, manual workaround or resource constraint threatens the service. Leaders assign owners and deadlines, accept residual exposure only through the bank’s defined process, and direct investment toward the weaknesses that most affect customer and business outcomes.

Metrics track whether the service remains within tolerance, whether dependencies and tests are current and whether remediation is effective. After a disruption, leaders preserve accountability while learning from the event, updating the map, assumptions, controls and scenarios instead of treating restoration as the end of the work.

Sources

Read the primary material

Banking Explained prioritizes regulators, official publications and first-party announcements.