By David Honour
World models, AI systems that model how real-world environments change, are expected to be one of the next steps in AI developments. These could help organizations explore disruption, test resilience choices, and rehearse responses. Some capabilities already exist, but the challenge is turning them into dependable resilience tools – and remaining capable when the model is unavailable, wrong, or manipulated.
An incident manager considering a move to a backup site needs more than a list of recovery options. Is the site actually accessible? Can its systems handle the additional workload? Does it depend on the same failing infrastructure? What happens to customers while the move takes place? World models could help organizations investigate such connected questions before committing to an intervention. Potentially, they provide a working representation of a physical environment through which possible changes and responses can be explored.
Simulation, scenario analysis, and digital twins already serve parts of this purpose. The opportunity is to extend their reach: making models easier to develop and adapt, exploring more combinations of events, and connecting operational evidence more closely.
What are world models?
In AI, a world model is generally a representation used to predict how an environment may evolve, often in response to possible actions. Many current approaches learn those dynamics from data and experience.
The ‘world’ need not mean the planet. It could be a robot’s workspace, a website, a warehouse, or a transport network. A model might represent stock movements, available capacity, and the effects of changing a dispatch schedule. It need not recreate every detail, but it must retain what matters for its intended use.
World representations can take the form of images, generated video, structured information, text, or internal mathematical representations that people cannot directly inspect. Language models can be components of world-model systems; this is not a simple division between machines that generate words and machines that understand reality.
Prediction, planning, and execution are different functions. A world model predicts possible changes. A planner or decision-maker uses predictions alongside objectives and constraints to select actions. A person or authorised system executes them.
Is this essentially a digital twin?
There is an important connection, but world models and digital twins are overlapping concepts rather than interchangeable labels.
A digital twin represents particular real-world entities or processes, with synchronisation at an appropriate frequency and level of detail. Digital twins already support prediction, simulation, and decision-making; those capabilities are not newly introduced by world models. A learned world model could supply some of the predictive dynamics within a twin.
For resilience, useful progress may look more like interconnected, separately governed models: for example a factory model linked to logistics, energy, and weather models, with agreed interfaces and explicit limits on what each represents.
Seven potential organizational resilience advantages
A living representation of the impact landscape
Where business impact analysis produces a periodic snapshot, a world model connected to appropriate operational data could help maintain a more current representation of critical services and their dependencies.
Changes in staffing, stock, transaction volumes, or supplier capacity could trigger reassessment of how disruption might propagate and whether response arrangements remain viable. The benefit would be a closer connection between the impact landscape and the organization’s changing operating conditions.
The model would need to distinguish verified dependencies from inferred ones and show when information was last checked. It could inform assessments of potential harm, but responsibility for defining unacceptable outcomes and approving impact tolerances would remain with accountable people.
Exploring compound and cascading disruption
World-model-enabled simulations could make it easier to explore combinations and sequences: a cloud outage during flooding, a cyber attack alongside workforce absence, or a port closure followed by a supplier failure.
The aim would be to find recurring vulnerabilities and test interventions across varied assumptions. If many scenarios exhaust the same specialist team or overwhelm the same recovery dependency, that is a reason to investigate the underlying constraint.
Scenario variety must not be confused with probability. Running thousands of simulations does not establish how likely their outcomes are.
For resilience work, uncalibrated outputs should support exploration of possibilities, not be presented as measured real-world probabilities.
Resilience by design
Models could help teams test the resilience implications of proposed changes before implementation: outsourcing, technology redesign, facility consolidation, staffing changes, or inventory policies.
Alternative designs could be exposed to the same disruptions to examine trade-offs among cost, service performance, recovery capability, and stakeholder harm. This could reveal that apparent redundancy shares a critical dependency, or that an efficiency improvement removes the capacity needed during disruption.
Modelling service degradation, rather than simple downtime, is a particularly relevant area. Resilience teams could explore which functions to reduce or suspend while preserving critical outcomes and whether the proposed reduced service remains workable for those who depend on it.
More frequent exercising and capability development
Generative environments could support exercises with changing information, resource constraints, and consequences that respond to participants’ decisions. Teams could repeat an exercise under different assumptions or compare approaches to the same disruption.
The distinction between exercising a model and exercising an organization is essential. Thousands of automated runs do not demonstrate that people can coordinate, challenge information, or execute recovery actions under the pressures of a real situation.
Human-led exercises would still need to introduce surprises outside the model’s assumptions. Rehearsal should build familiarity with response capabilities while testing adaptability – not reward participants simply for learning what the simulation expects.
Dynamic incident decision support
During an incident, a validated model could help compare interventions against the current situation: diverting work, switching suppliers, isolating systems, reallocating staff, or operating a reduced service.
The model would need to make uncertainty and unavailable information visible. It should not silently treat an unconfirmed resource as available, or assume that two response modules can use the same people simultaneously. Initial operational use would need to support accountable decision-makers; authority to execute actions requires a separate justification and explicit limits.
Visibility of systemic and cross-sector risk
Federated models could help organizations examine dependencies that cross their boundaries, including energy, telecommunications, cloud services, finance, and logistics. They could support coordinated exercises and explore where an intervention might reduce wider disruption.
Faster learning and adaptation
Comparing predicted and observed outcomes could help expose missing dependencies, faulty assumptions, and informal workarounds. Those findings could inform both model revision and changes to the organization’s resilience arrangements.
The practical opportunity is to make assumptions more inspectable and learning more durable. The corresponding requirement is to preserve evidence, challenge stale beliefs, and test whether learning remains valid when systems or providers change.
What might this look like in practice?
Consider an example of a regional distributor supplying food to care homes. Flooding restricts access to its main depot, a cloud outage disables its routing application, and staff absence reduces dispatch capacity.
A useful simulation would need more than a map. It would require verified information about stock, vehicles, staff, receiving sites, road access, and technology dependencies. Operational rules would constrain what can actually be moved, by whom, and when.
Using a world model, the resilience team could compare transferring work to another depot, using a pre-arranged manual dispatch process, engaging another carrier, or prioritising sites with the least remaining stock. The model might help estimate particular delays or capacity effects, provided those estimates had been validated for the conditions concerned.
The model might reveal that both depots depend on an identity service affected by the cloud outage. It might also expose uncertainty about another carrier’s available capacity.
Constraints and validation
A world model is a selective representation of reality. The states that matter most during disruption – hidden dependencies, damaged capacity, human availability, or adversarial activity – may be precisely those it cannot observe reliably. Validation would therefore be essential.
Before relying on a model, an organization would need to consider the following areas and take the suggested actions:
- Fit for which decision? Define the intended use, relevant time horizon, operating conditions, and unacceptable consequences.
- What is known? Identify observed, inferred, and missing states, data freshness, and constraints the simulation must respect.
- What has been tested? Use independent evidence, historical incidents not used in development, practical tests, and deliberately challenging compound scenarios. Assess intermediate states as well as final outputs.
- When must it be challenged or withdrawn? Establish triggers for rechecking data, limiting recommendations, or suspending use when uncertainty or prediction errors increase.
- Who can act? Separate recommendation from execution, constrain permissions, require appropriate approval, and provide ways to halt or reverse actions where feasible.
- Can it be trusted and replaced? Protect data and retained beliefs against manipulation; maintain provenance, version control, and rollback; test provider substitution and operation without the model.
Replaying past incidents provides useful evidence, but cannot prove performance in an unprecedented crisis. Confidence should be tied to validation for the intended use, not simply to the confidence expressed by an AI system.
Multiple models also do not automatically provide independent assurance. They may share training data, underlying models, cloud infrastructure, or identity services.
Conclusion
World models have strong potential for future organizational resilience. They may be able to help people ask better questions about connected systems, make better continuity, response, and recovery plans, test and exercise scenarios more comprehensively, and better manage real incidents. Operationally, they aren’t here yet; but blink and they might be!
The author
David Honour is editor of Resilience Forward.
Both Chat GPT-6 Astra (Max) and Gemini 3.1 Pro helped provide and assess the background information for this article.






