Introduction
Testing is crucial to assess an organization’s impact tolerances and determine whether its incident response strategies will enable the organization to recover a business service within the defined impact tolerance. Testing also gives the organization a clear understanding of the severe but plausible scenarios that might create maximum stress to business services. Severe but plausible scenarios refer to situations that would “result in high impact and significant disruptions and, while unlikely to occur, remain probable,” (Alexander 2023).
Key points about scenario testing include:
- Organizations should ensure that their approach to testing and determining scenarios has received appropriate challenge from senior management and receives their endorsement.
- Organizations will need to have a framework in place outlining their approach to conducting scenario tests. In particular, organizations will need to consider how they have determined their criteria for formulating severe but plausible service-specific tests (including factors that may complicate an organization’s recovery). As part of their testing methodology, organizations should also consider the availability of workarounds and substitutes.
- Organizations should test a range of scenarios, including those in which they anticipate exceeding their impact tolerance. Understanding the circumstances where it is impossible to stay within an impact tolerance will provide useful information to the organization. Boards and senior management will need to judge whether failing to remain within the impact tolerance in specific scenarios is acceptable and will need to be able to explain their reasoning to regulators.
- Exercises are a common method of testing severe but plausible scenarios. An exercise program should combine both announced and unannounced exercise methods. One of the main shortcomings of announced exercises is that they fail to test the state of constant readiness and the ability of the business continuity and crisis management teams to react quickly. The unannounced exercise method is very important and is not used as often as it should be in many organizations. Incidents occur without warning. Companies should put more emphasis on unannounced exercises for testing the constant readiness of the teams and their ability to react to the surprise announcement of the exercise.
A scenario testing methodology
Regulators expect organizations to develop a scenario testing plan that details how they will gain assurance that they can remain within impact tolerances for important business services.
For exercises to be effective they have to be well planned. A methodology is presented below that details the sequence of steps that should be followed to plan and execute a scenario testing exercise.

1) Review previous scenario testing exercises
When planning a future event, the results of previous scenario testing exercises should always be considered. There should always be a natural progression from simple to complex exercises.
The information from previous exercises is very useful in identifying the current exercise objectives. Some of the most important issues in reviewing past exercises help determine:
- Components and areas of the scenario testing plan that have not been part of an exercise.
- Components and processes of the scenario testing plan that did not work.
Challenges and obstacles encountered during execution of the previous scenario testing exercises can help identify exercise risks in the current plan.
2) Identifying severe but plausible scenarios
A plausible scenario is a realistic event that could disrupt the delivery of important or critical business services, leading to unacceptable impacts.
When setting scenarios, organizations could consider previous incidents or near misses within the organization, across their sector, and in other sectors and jurisdictions.
A scenario library
At this point a scenario library is strongly recommended. Having a scenario library provides a framework through which business service-specific scenarios can be developed and relevant complicating factors added to ensure that they meet the necessary ‘severe but plausible’ criteria.
A scenario library acts as a repository for generic, real-life, severe but plausible scenarios that can be used to design business service-specific scenario tests. The scenarios contained within this library benefit from taking into consideration the known risks and threats affecting the organization’s important business services, the organization itself, and the wider market. The scenario library should be regularly reviewed and updated to reflect the latest risks identified from the organization’s broader horizon scanning and intelligence gathering activities.
3) Identify exercise objectives and scope
Exercise objectives define the criteria for success. It is essential to define the exercise objectives in very precise and measurable terms. Normally the time allocated and the budget available for exercises is limited. The main objective of scenario testing is to test and exercise the organization’s ability to maintain its important business services within the approved impact tolerances.
The exercise scope, which identifies the overall depth and breadth of the exercise, is to test the ability of the organization to validate if an important business service can be maintained within the approved impact tolerances.
4) Assess exercise constraints
The “exercise constraints are elements that limit or restrict the options available for conducting the exercise,” (Alexander 2016). A clear understanding of the exercise constraints and their potential effects on an exercise is essential for developing a viable test strategy, logistics, and schedule. Some examples of possible exercise constraints could be:
- Financial constraints: a limited exercise budget due to financial constraints can affect the exercise in different ways.
- Security restrictions: the exercise may require access to confidential data and transactions and sensitive systems and facilities.
- Availability of recovery test teams: the availability of team members can become a constraint. Team members may have planned vacations or may have vital day to day commitments during the test period.
5) Design the exercise strategy
This step defines a strategy to achieve the test objectives. The strategy information related to any current test objectives in previous exercise plans can be used in this step as a basis for developing the current test strategy. One of the purposes of this step in the proposed methodology is to ensure that the exercise plan is consistent with the constraints identified in step four.
An exercise strategy has different components that need to be considered in the development of the test strategy. These are described below:
Timing: establishing a date, time, and duration of the exercise requires careful consideration of the constraints and the availability of required resources. Exercise timing is determined through an evaluation of the readiness and availability of various resources, such as: test software and data; specialized test equipment; business recovery exercise teams; and recovery hardware.
As a general rule, the timing should minimize impacts to normal business operations and avoid peak workloads, holidays, and important business events.
Exercise and testing methods: the choice of exercise method is informed by the degree of maturity of an organization’s operational program. Examples include:
Drill – these tend to be informal, team-based exercises that test scenarios at individual asset level. The main objective is to walk through roles and responsibilities.
Desktop exercises or workshops – these are typically structured as walkthroughs of a recovery plan. They are usually less resource intensive and would typically involve management and, where feasible, third party providers.
Internal simulated tests – these are facilitated, simulated, tests involving both business and functional stakeholders to test the response to a severe but plausible disruption.
External simulated tests – these require organizations to coordinate testing with material third party providers which support their important business services.
Full live testing – such testing would typically be conducted in the production environment and entail creating a real-time disruption to test the organization’s ability to remain within impact tolerances.
Live parallel testing – this does not disrupt the organization’s production systems but rather tests its failover arrangements and is less risky than full live testing.
Incidents – real incidents and near misses can also help organizations assess the effectiveness of their resilience measures.
Third party assurance – organizations may also consider using alternative forms of assurance from third parties, such as previous disaster recovery test results, independent reports, and third-party certifications.
As a guideline, when planning scenario testing organizations should consider the following key features, that were presented by Duncan Mackinnon, of the Prudential Regulation Authority (PRA), which set out where firms in financial services in the UK are expected to focus as they work towards building operational resilience by March 2025:
- Assume disruption has occurred,
- Include data integrity issues,
- Incorporate third party disruption,
- Consider factors beyond the firm’s control,
- Ask what might happen if backup arrangements do not function as anticipated,
- Include cases where multiple parts of the organization are disrupted simultaneously, (Aldbury International 2021).
As a guideline, simpler and basic scenario testing should be carried out prior to more complex tests.
6) Exercise logistics
The logistics process plays a very important role in testing scenarios. Exercise logistics is a process that deals mainly with four areas of concern:
- Formation of operational resilience test teams: the size, structure, and members of the team depend on the test objectives and scope.
- Test resource procurement: to be able to ensure timely availability of required resources, a detailed list of resource procurement tasks is prepared and executed well in advance of a test date. Advance preparations are a key to minimizing costs and impacts to test timing due to unexpected delays in resource order processing, shipment, and setup.
- Mobilization of personnel: a test may require mobilization of business operational resilience test teams to remote locations such as off-site storage facilities, alternate IT recovery facilities, alternate manufacturing and production facilities, alternate office work areas, and the crisis management facility. Planning and implementing logistics activities are crucial to mobilize operational resilience test teams.
- Test facilities provisioning: the test plan should include logistics activities to ensure the availability of test facilities that can adequately support the test requirements.
7) Development of exercise schedule
An exercise or test schedule, like any other project schedule, demands careful planning and management skills. A “test schedule details the list of recovery activities, procedures, tasks, priorities, assignments, start and end dates and times, and dependencies,” (Alexander 2016). Usually a test schedule divides activities into three phases:
- Test preparation phase,
- Test execution phase, and
- Test evaluation phase.
The test preparation phase begins once the operational resilience test plan document is developed. The activities in this phase include pre-test meetings with offsite storage vendors and alternate recovery facilities providers. The logistics activities of the test plan begin during the test preparation phase.
The actual test is conducted during the test execution phase at the date and time specified by the test strategy. This phase usually covers recovery activities that are part of the test objectives. The test schedule organizes these activities into appropriate recovery strategies.
The test evaluation phase begins immediately after the completion of the test execution phase. The test schedule should include activities to evaluate the test results, produce an evaluation report, and present the contents of the evaluation report to management. The main purpose of these activities is to evaluate the extent of success in achieving the test objectives, teams’ performance, problems encountered during the test, and gaps and weaknesses observed in the business continuity and crisis management arrangements. An important part of the test evaluation phase is to validate whether the organization’s incident response playbook is fit for purpose and that it will enable the organization to recover the service within the tolerance defined.
8) Exercise risks identification
This step of the methodology identifies and controls potential risks of tests failures based on a thorough review of all the information gathered in the preceding steps.
Once the risks are identified, the teams should review the risks, determine possible solutions for minimizing the risks, and incorporate the accepted solutions into the operational resilience test plan.
9) Post-exercise written report
Once the exercise is over the organization will have invested considerable resources in its design and delivery. The organization needs to develop a process that will generate information to assess the effectiveness of the exercise and which will allow the organization to initiate improvement actions.
A comprehensive report is required, which should reflect what took place on the day and be linked clearly to the exercise aims and objectives. Once the report is completed, it should provide objective evidence for:
- Identifying amendments to the existing incident management plan, supporting procedures, and processes. Alternatively identifying requirements for producing a new plan.
- Identifying and justifying future training requirements for individuals and teams.
- Identifying and justifying additional resources to enhance the current capability.
- Identifying objectives for future exercises.
- Providing audit evidence of the effectiveness of the company’s approach to incident management.
Some guidelines for scenario testing in operational resilience
Here are some guidelines to be considered when testing scenarios in operational resilience:
- The board plays a pivotal role in operational resilience, particularly in scenarios where regulatory discussions are at the forefront. Central banks or regulatory authorities may scrutinize the organization’s resilience strategies; and the board must be prepared to engage in meaningful dialogue.
- Scenario testing is more than an exercise; it is an opportunity to evaluate how the organization responds and recovers from severe disruptions. Unlike traditional business continuity or crisis management testing, operational resilience scenario testing extends beyond the internal scope. It necessitates the involvement of third-party entities and customers, adding complexity to the testing landscape.
- It is essential to prove that the scenarios are not just another replication of traditional business continuity exercises but also scenarios that can test impact tolerances and readiness.
- Keep an eye on industry peers and competitors. Monitor their scenario-testing approaches and learn from their experiences.
- Lean on the extreme but plausible side when crafting scenarios.
- Emphasize the importance of feedback, review, and debriefing after each test.
- Operational resilience is a dynamic field requiring organizations to evolve their strategies and practices continuously.
- Meeting regulatory expectations in scenario testing demands a thoughtful approach, prioritizing prevention, active involvement, and adaptability.
- By balancing severe scenarios and improbable events, organizations can demonstrate their commitment to resilience while effectively navigating the complexities of regulatory compliance.
Conclusions
Scenario testing is a key part of operational resilience. It builds and demonstrates capability to respond and recover within pre-defined impact tolerance levels.
A range of extreme, yet plausible, shock scenarios that impact the resources required to deliver the important business service that is being tested has to be identified. These may include such things as natural disasters, pandemics, social unrest, conflict, cyber attacks, or infrastructure issues. These scenarios need to be tested to find out if the impact tolerances would be met.
The only way an organization can evaluate its ability to remain within impact tolerances in severe but plausible disruption scenarios for each of its business services is through a well-developed scenario testing methodology.
Scenario testing enables firms to gain a comprehensive understanding of the resilience of their important business services and to identify areas where action needs to be taken to remediate vulnerabilities to build resilience over time.
Understanding the severe but plausible scenarios where a organization is unable to remain within the impact tolerances that it has set is just as important as understanding the instances in which an organization can meet its tolerances.
The Board may also need to be engaged to determine whether additional investment is needed to address findings from scenarios where organizations would breach their impact tolerances.
The Board owns the operational resilience approach in the organization and plays a pivotal role.
Scenario testing identifies weaknesses, exposes vulnerabilities, and provides insights that mere theoretical planning cannot.
By simulating real-world scenarios, an organization gains a first-hand understanding of how it will perform in times of crisis.
Bibliographical references
- Alexander, Alberto. Planning and Managing Business Continuity Management Arrangements, Continuity Central, 2016.
- Alexander, Alberto Managing Operational Resilience, Resilience Forward, 2023.
- Alexander, Alberto, Developing and Managing Impact Tolerances, Resilience Forward, 2024.
- Aldbury International Limited, London, UK, 2020.
The author
Alberto G. Alexander, Ph.D, MBCI, International Consultant
Dr. Alexander holds a Ph.D from The University of Kansas and a M.A from Northern Michigan University. He is the Managing Director of the international consulting and training firm Eficiencia Gerencial y Productividad, located in Lima, Perú. He is a Member of the Business Continuity Institute and can be contacted at alexander@egpsac.com
He is currently Professor at the Graduate Business School of ESAN University, Lima, Peru.






