0% found this document useful (0 votes)
18 views7 pages

Understanding Complex System Failures

The document discusses the nature of failure in complex systems, emphasizing that these systems are inherently hazardous and require multiple layers of defense against failure. It argues that catastrophic failures arise from the combination of multiple small failures rather than a single point of failure, and that human practitioners play a dual role in both production and defense against failure. Additionally, it highlights the importance of understanding that safety is an emergent property of the system as a whole, rather than a characteristic of individual components.

Uploaded by

João Marcos Q
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views7 pages

Understanding Complex System Failures

The document discusses the nature of failure in complex systems, emphasizing that these systems are inherently hazardous and require multiple layers of defense against failure. It argues that catastrophic failures arise from the combination of multiple small failures rather than a single point of failure, and that human practitioners play a dual role in both production and defense against failure. Additionally, it highlights the importance of understanding that safety is an emergent property of the system as a whole, rather than a characteristic of individual components.

Uploaded by

João Marcos Q
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

5/8/25, 10:50 AM How Complex Systems Fail

How Complex Systems Fail


(Being a Short Treatise on the Nature of Failure; How Failure is Evaluated; How Failure is
Attributed to Proximate Cause; and the Resulting New Understanding of Patient Safety)
Richard I. Cook, MD
Cognitive Technologies Labratory
University of Chicago
1. Complex systems are intrinsically hazardous systems.
All of the interesting systems (e.g. transportation, healthcare, power
generation) are inherently and unavoidably hazardous by the own nature. The
frequency of hazard exposure can sometimes be changed but the processes
involved in the system are themselves intrinsically and irreducibly hazardous. It
is the presence of these hazards that drives the creation of defenses against
hazard that characterize these systems.
2. Complex systems are heavily and successfully defended against failure
The high consequences of failure lead over time to the construction of multiple
layers of defense against failure. These defenses include obvious technical
components (e.g. backup systems, ‘safety’ features of equipment) and human
components (e.g. training, knowledge) but also a variety of organizational,
institutional, and regulatory defenses (e.g. policies and procedures,
certification, work rules, team training). The effect of these measures is to
provide a series of shields that normally divert operations away from accidents.
3. Catastrophe requires multiple failures – single point failures are not
enough.
The array of defenses works. System operations are generally successful.
Overt catastrophic failure occurs when small, apparently innocuous failures join
to create opportunity for a systemic accident. Each of these small failures is
necessary to cause catastrophe but only the combination is sufficient to permit
failure. Put another way, there are many more failure opportunities than overt
system accidents. Most initial failure trajectories are blocked by designed
system safety components. Trajectories that reach the operational level are
mostly blocked, usually by practitioners.
[Link] 1/7
4. Complex systems contain changing mixtures of failures latent within
5/8/25, 10:50 AM How Complex Systems Fail

them.
The complexity of these systems makes it impossible for them to run without
multiple flaws being present. Because these are individually insufficient to
cause failure they are regarded as minor factors during operations. Eradication
of all latent failures is limited primarily by economic cost but also because it is
difficult before the fact to see how such failures might contribute to an
accident. The failures change constantly because of changing technology,
work organization, and efforts to eradicate failures.
5. Complex systems run in degraded mode.
A corollary to the preceding point is that complex systems run as broken
systems. The system continues to function because it contains so many
redundancies and because people can make it function, despite the presence
of many flaws. After accident reviews nearly always note that the system has a
history of prior ‘proto-accidents’ that nearly generated catastrophe. Arguments
that these degraded conditions should have been recognized before the overt
accident are usually predicated on naïve notions of system performance.
System operations are dynamic, with components (organizational, human,
technical) failing and being replaced continuously.
6. Catastrophe is always just around the corner.
Complex systems possess potential for catastrophic failure. Human
practitioners are nearly always in close physical and temporal proximity to
these potential failures – disaster can occur at any time and in nearly any place.
The potential for catastrophic outcome is a hallmark of complex systems. It is
impossible to eliminate the potential for such catastrophic failure; the potential
for such failure is always present by the system’s own nature.
7. Post-accident attribution to a ‘root cause’ is fundamentally wrong.
Because overt failure requires multiple faults, there is no isolated ‘cause’ of an
accident. There are multiple contributors to accidents. Each of these is
necessarily insufficient in itself to create an accident. Only jointly are these
causes sufficient to create an accident. Indeed, it is the linking of these causes
together that creates the circumstances required for the accident. Thus, no
isolation of the ‘root cause’ of an accident is possible. The evaluations based
[Link] 2/7
on such reasoning as ‘root cause’ do not reflect a technical understanding of
5/8/25, 10:50 AM How Complex Systems Fail

the nature of failure but rather the social, cultural need to blame specific,
localized forces or events for outcomes. 1
1 Anthropological field research provides the clearest demonstration of the social
construction of the notion of ‘cause’ (cf. Goldman L (1993), The Culture of
Coincidence: accident and absolute liability in Huli, New York: Clarendon Press; and
also Tasca L (1990), The Social Construction of Human Error, Unpublished doctoral
dissertation, Department of Sociology, State University of New York at Stonybrook)

8. Hindsight biases post-accident assessments of human performance.


Knowledge of the outcome makes it seem that events leading to the outcome
should have appeared more salient to practitioners at the time than was
actually the case. This means that ex post facto accident analysis of human
performance is inaccurate. The outcome knowledge poisons the ability of
after-accident observers to recreate the view of practitioners before the
accident of those same factors. It seems that practitioners “should have
known” that the factors would “inevitably” lead to an accident. 2 Hindsight bias
remains the primary obstacle to accident investigation, especially when expert
human performance is involved.
2 This is not a feature of medical judgements or technical ones, but rather of all
human cognition about past events and their causes.

9. Human operators have dual roles: as producers & as defenders against


failure.
The system practitioners operate the system in order to produce its desired
product and also work to forestall accidents. This dynamic quality of system
operation, the balancing of demands for production against the possibility of
incipient failure is unavoidable. Outsiders rarely acknowledge the duality of this
role. In non-accident filled times, the production role is emphasized. After
accidents, the defense against failure role is emphasized. At either time, the
outsider’s view misapprehends the operator’s constant, simultaneous
engagement with both roles.
0. All practitioner actions are gambles.
[Link] 3/7
After accidents, the overt failure often appears to have been inevitable and the
5/8/25, 10:50 AM How Complex Systems Fail

practitioner’s actions as blunders or deliberate willful disregard of certain


impending failure. But all practitioner actions are actually gambles, that is, acts
that take place in the face of uncertain outcomes. The degree of uncertainty
may change from moment to moment. That practitioner actions are gambles
appears clear after accidents; in general, post hoc analysis regards these
gambles as poor ones. But the converse: that successful outcomes are also
the result of gambles; is not widely appreciated.
1. Actions at the sharp end resolve all ambiguity.
Organizations are ambiguous, often intentionally, about the relationship
between production targets, efficient use of resources, economy and costs of
operations, and acceptable risks of low and high consequence accidents. All
ambiguity is resolved by actions of practitioners at the sharp end of the
system. After an accident, practitioner actions may be regarded as ‘errors’ or
‘violations’ but these evaluations are heavily biased by hindsight and ignore the
other driving forces, especially production pressure.
2. Human practitioners are the adaptable element of complex systems.
Practitioners and first line management actively adapt the system to maximize
production and minimize accidents. These adaptations often occur on a
moment by moment basis. Some of these adaptations include: (1)
Restructuring the system in order to reduce exposure of vulnerable parts to
failure. (2) Concentrating critical resources in areas of expected high demand.
(3) Providing pathways for retreat or recovery from expected and unexpected
faults. (4) Establishing means for early detection of changed system
performance in order to allow graceful cutbacks in production or other means
of increasing resiliency.
3. Human expertise in complex systems is constantly changing
Complex systems require substantial human expertise in their operation and
management. This expertise changes in character as technology changes but
it also changes because of the need to replace experts who leave. In every
case, training and refinement of skill and expertise is one part of the function of
the system itself. At any moment, therefore, a given complex system will
contain practitioners and trainees with varying degrees of expertise. Critical
[Link] 4/7
issues related to expertise arise from (1) the need to use scarce expertise as a
5/8/25, 10:50 AM How Complex Systems Fail

resource for the most difficult or demanding production needs and (2) the
need to develop expertise for future use.
4. Change introduces new forms of failure.
The low rate of overt accidents in reliable systems may encourage changes,
especially the use of new technology, to decrease the number of low
consequence but high frequency failures. These changes maybe actually
create opportunities for new, low frequency but high consequence failures.
When new technologies are used to eliminate well understood system failures
or to gain high precision performance they often introduce new pathways to
large scale, catastrophic failures. Not uncommonly, these new, rare
catastrophes have even greater impact than those eliminated by the new
technology. These new forms of failure are difficult to see before the fact;
attention is paid mostly to the putative beneficial characteristics of the
changes. Because these new, high consequence accidents occur at a low rate,
multiple system changes may occur before an accident, making it hard to see
the contribution of technology to the failure.
5. Views of ‘cause’ limit the effectiveness of defenses against future events.
Post-accident remedies for “human error” are usually predicated on
obstructing activities that can “cause” accidents. These end-of-the-chain
measures do little to reduce the likelihood of further accidents. In fact that
likelihood of an identical accident is already extraordinarily low because the
pattern of latent failures changes constantly. Instead of increasing safety, post-
accident remedies usually increase the coupling and complexity of the system.
This increases the potential number of latent failures and also makes the
detection and blocking of accident trajectories more difficult.
6. Safety is a characteristic of systems and not of their components
Safety is an emergent property of systems; it does not reside in a person,
device or department of an organization or system. Safety cannot be
purchased or manufactured; it is not a feature that is separate from the other
components of the system. This means that safety cannot be manipulated like
a feedstock or raw material. The state of safety in any system is always
[Link] 5/7
dynamic; continuous systemic change insures that hazard and its management
5/8/25, 10:50 AM How Complex Systems Fail

are constantly changing.


7. People continuously create safety.
Failure free operations are the result of activities of people who work to keep
the system within the boundaries of tolerable performance. These activities
are, for the most part, part of normal operations and superficially
straightforward. But because system operations are never trouble free, human
practitioner adaptations to changing conditions actually create safety from
moment to moment. These adaptations often amount to just the selection of a
well-rehearsed routine from a store of available responses; sometimes,
however, the adaptations are novel combinations or de novo creations of new
approaches.
8. Failure free operations require experience with failure.
Recognizing hazard and successfully manipulating system operations to
remain inside the tolerable performance boundaries requires intimate contact
with failure. More robust system performance is likely to arise in systems where
operators can discern the “edge of the envelope”. This is where system
performance begins to deteriorate, becomes difficult to predict, or cannot be
readily recovered. In intrinsically hazardous systems, operators are expected to
encounter and appreciate hazards in ways that lead to overall performance
that is desirable. Improved safety depends on providing operators with
calibrated views of the hazards. It also depends on providing calibration about
how their actions move system performance towards or away from the edge of
the envelope.
Other Materials
Cook, Render, Woods (2000). Gaps in the continuity of care and progress on
patient safety. British Medical Journal 320: 791-4.
Cook (1999). A Brief Look at the New Look in error, safety, and failure of
complex systems. (Chicago: CtL).
Woods & Cook (1999). Perspectives on Human Error: Hindsight Biases and
Local Rationality. In Durso, Nickerson, et al., eds., Handbook of Applied
Cognition. (New York: Wiley) pp. 141-171.
[Link] 6/7
Woods & Cook (1998). Characteristics of Patient Safety: Five Principles that
5/8/25, 10:50 AM How Complex Systems Fail

Underlie Productive Work. (Chicago: CtL)


Cook & Woods (1994), “Operating at the Sharp End: The Complexity of Human
Error,” in MS Bogner, ed., Human Error in Medicine , Hillsdale, NJ; pp. 255-
310
Woods, Johannesen, Cook, & Sarter (1994), Behind Human Error: Cognition,
Computers and Hindsight , Wright Patterson AFB: CSERIAC.
Cook, Woods, & Miller (1998), A Tale of Two Stories: Contrasting Views of
Patient Safety, Chicago, IL: NPSF
Original Copyright © 1998, 1999, 2000 by [Link], MD, for CtL

[Link] 7/7

Common questions

Powered by AI

Practitioners create safety in complex systems through continuous adaptation to changing conditions, ensuring that operations stay within tolerable performance limits . This process involves selecting appropriate routines or devising novel approaches to manage emerging risks. Practitioners also adapt by restructuring systems, concentrating critical resources, and establishing recovery pathways to enhance resilience . Such adaptations are essential as complex systems are inherently dynamic, and safety emerges as an ongoing activity from these human interventions .

Human practitioners ensure the dynamic operation of complex systems by simultaneously engaging in production and safeguarding against system failures . Before an accident, the emphasis is on their role in maintaining productivity. However, after an accident, the focus shifts to their responsibility for preventing failures . This change in perception often leads to an imbalance in evaluating their roles, not accounting for the dual responsibilities they continuously manage to keep the system functioning effectively in both scenarios .

Complex systems are intrinsically hazardous due to the nature of the processes involved, making them inherently and irreducibly hazardous . To mitigate the risk of accidents, these systems are equipped with multiple layers of defenses. These include technical components like backup systems and safety features, along with human, organizational, institutional, and regulatory defenses . While these defenses generally prevent accidents, catastrophic events occur when small, innocuous failures combine to bypass these shields . Thus, the defenses play a crucial role in normal operations by blocking potential failure trajectories .

The integration of novel technology in complex systems impacts safety by potentially introducing new forms of failure, as the new technologies may create unforeseen pathways to high consequence accidents . While aiming to improve precision and reduce known issues, these technologies often lead to rare but severe failures. Precautionary measures should involve thorough foresight and evaluation processes to assess potential failure modes before implementation. Additionally, building robust monitoring mechanisms that signal deviations from expected performance can help mitigate risks posed by these technological changes .

Hindsight bias significantly hinders post-accident assessments by making it appear that the outcomes and the events leading up to them should have been more evident to practitioners at the time . The knowledge of the outcome influences observers to overestimate the predictability of the accident in hindsight, skewing evaluations of human performance. This cognitive bias makes post hoc analysis inaccurate because it distorts the practitioners' real-time perspective by suggesting they should have anticipated the failure . This bias remains a primary hurdle in accurately interpreting human actions and decisions pre-accident .

Identifying a 'root cause' for failures in complex systems is flawed because accidents result from multiple interconnected faults rather than a single isolated cause . These faults, individually insufficient to cause a failure, become significant only when they interact and accumulate, thus forming a systemic failure . The concept of a 'root cause' falsely simplifies the complexity of the interactions and circumstances that lead to an accident, disregarding the multifactorial nature of failures in such systems .

Human operators in complex systems perform dual roles: they must both produce the desired output of the system and act as a defense against potential failures . This dynamic involves continuously balancing production demands with guarding against incipient system failures. Outsiders often misunderstand this duality by emphasizing one role over the other depending on whether there is an accident or not, thus failing to appreciate the simultaneous engagement required for both roles . During non-accident times, outsiders focus more on the production aspect, whereas, after accidents, they emphasize the defensive role .

Safety in complex systems is an emergent property that arises from the interaction and integration of all system components rather than any single element . It cannot be isolated or purchased like a standalone feature because it dynamically evolves with continuous systemic changes involving people, technology, and organizational structure. As safety results from how these parts collectively function and adapt, constant management and adjustments are required to maintain an acceptable level of performance .

Experience with failure is crucial for failure-free operations because it helps operators recognize the bounds of safe system performance and take necessary actions to avoid crossing these limits . Familiarity with failure enables practitioners to discern when systems begin to deteriorate, leading to proactive measures that maintain operations within safe parameters. Such experience provides operators with calibrated views of hazards, enhancing their ability to navigate the system safely and increase overall resiliency .

Changes in complex systems, such as adopting new technology, might aim to reduce low consequence, high-frequency failures but often introduce opportunities for new, high consequence failures . These changes are complex, as each new technology might create unpredictable pathways to catastrophic failures, offering a different failure mode from the one it intended to eliminate . The challenge arises because these failures are difficult to foresee before implementation, and by the time a new catastrophic failure occurs, multiple system changes may have obfuscated the role of the new technology .

You might also like