Even though the adage ‘fail to prepare, prepare to fail’ is 200 years old, it could just as easily have been invented 200 days ago as a call to action for contemporary technology and business resilience. According to research published last year, for example, only 20% of organizations say they are fully prepared for an unplanned disruption to a system or service. In defence of the other 80%, implementing the infrastructure and processes required to prevent and mitigate technology outages can be extremely challenging, especially at scale. Yet the consequences of a resilience failure can be so severe that many organizations need to reassess their risk exposure and capacity to recover.
Cyber attacks have been strongly in focus over the past few years, driven by high-profile incidents. But herein lies an important problem: the widespread tendency to focus on security as the primary risk to resilience, commonly at the expense of everything else. In reality, some of the most serious technology resilience incidents have other root causes, with infrastructure failure, software issues, misconfiguration, or third-party dependencies among them. Add the dramatic effect of global AI implementation to the mix, and scaling resilience to cover all internal and external dependencies is more complex than ever.
Resilience is an IT design principle
Resilience, therefore, should not be defined by the absence of failure or even by trying to ensure 100% protection 100% of the time, but by the ability to maintain operations if, and when, a failure occurs.
For today’s highly distributed organizations, visibility across the entire technology environment is critical and must include all of the various in-house resources, third-party service providers, and everyone in between. When disruption happens – and whatever the root cause – maintaining operations fundamentally depends on how systems are designed.
Without appropriate levels of insight across all infrastructure components, issues may appear resolved at one layer while continuing to impact downstream systems and users. Resilient architectures must also recover quickly, ensuring that services can be restored without prolonged disruption. This includes designing workflows that can withstand partial failure, rather than assuming continuous availability.
In practice, this means building systems that can continue operating in a degraded state, rather than falling over completely. At scale, resilience is therefore less about avoiding disruption and more about maintaining continuity despite it.
A layered approach
While system design offers a strong foundation for resilience, good outcomes are determined as much by decision-making and coordination as anything else. In particular, these capabilities depend on building well-defined, rehearsed incident response processes rather than relying on ad hoc decisions made under the most extreme pressure and time constraints.
For all the stakeholders involved in responding to an outage, including board-level decision-makers and beyond, roles and responsibilities must be clearly understood, so that specific teams can act quickly to expedite the appropriate response without ambiguity.
As anyone who has been at the sharp end of a technology infrastructure outage at scale will know, these situations can be very fluid. Everyone involved must have access to timely and accurate information, backed by effective communication and escalation processes, so that decisions are based on current system conditions. Without this, organizations risk making changes that address symptoms rather than underlying issues.
The data sovereignty blindspot
Another issue with scope to catch organizations off-guard is data sovereignty.
In basic terms, this is the requirement that data be stored, processed, and governed in accordance with the laws of the jurisdiction in which it resides, including who can access it and under what authority. As a result, it extends beyond actual physical location to include legal control and provider ownership, ensuring that organizations can demonstrate where data is held and which regulatory framework applies.
Given the market dominance of the large US-based cloud providers (particularly AWS, Microsoft, and Google), businesses are, to a greater or lesser extent, exposed to US data sovereignty rules and regulations. Many large organizations have outsourced significant and, in some cases, all of their IT and data to these third parties.
But is it a resilience priority? The answer is yes, absolutely. Consider this scenario: a large organization operating across Europe has critical data stored with service providers subject to external jurisdictions, such as those in place in the US. Following a major incident, access to that data is restricted or delayed due to regulatory controls or cross-border transfer limitations. While the systems may be technically recoverable, the organization is unable to access the information required to restore services. As a result, recovery is delayed and critical operations remain unavailable, with clear resilience implications.
The underlying point that draws all of the above challenges together is that resilience has become an increasingly complex and nuanced set of priorities and scenarios. Clearly, cyber security remains front and centre, but it’s far from the only reason why being able to continue operating under a wide range of conditions and facing disruption head on, has become a core organizational priority.
The author
Terry Storrar is Managing Director, Leaseweb UK






