Глава 1
When Risk Lurks Below the Surface: Mastering the Hidden Science of Reliability
In a world where catastrophes seem to strike without warning, K. Scott Griffith's groundbreaking approach to risk management offers a lifeline. His journey began on a routine day in 1985 when, as a young American Airlines pilot, he witnessed Delta Flight 191 plummet from the sky during a sudden microburst, killing most of the 163 people aboard. This traumatic event launched his decades-long quest to understand the patterns behind disasters and develop a science of reliability that could prevent such tragedies. Today, his methods have contributed to a remarkable 95% reduction in fatal airline accidents and transformed safety practices across healthcare, energy, and law enforcement. What makes this book particularly powerful is how it reveals the hidden patterns connecting seemingly unrelated catastrophes-from aviation disasters to medical errors to business failures-and offers a clear, scientific approach to managing the risks hiding in plain sight in every organization and everyday life.
Глава 2
The Iceberg of Risk: What Lies Beneath Our Successes
We interpret the world through our unique experiences, which sometimes helps us succeed but often creates dangerous blind spots. While we celebrate our successes, we sometimes learn the wrong lessons from them, paying with fortunes and lives when risks materialize. Business leaders typically focus on operational excellence rather than potential negative consequences, leaving organizations vulnerable to unexpected outcomes.
Consider Facebook's privacy failures or Apple's supply chain disruptions-both stemmed from hidden risks that weren't properly identified before they surfaced catastrophically. The "iceberg model" illustrates how visible problems represent just a fraction of dangers lurking beneath the surface. Organizations often develop false confidence from past successes, learning wrong lessons from good outcomes achieved through vulnerable systems until failure occurs.
This pattern appears across diverse catastrophes-from maritime disasters like the Titanic to space shuttle accidents, from business failures at companies like Enron to more recent challenges at Meta and Tesla. Our optimism typically outweighs our risk intelligence until we personally experience failure.
Effectively managing risk requires two distinct skills: seeing risk (having vision to detect what's ahead) and understanding risk (having intelligence to know how it affects us). Sometimes we understand dangers but can't see them because we aren't looking or they're beyond our field of view. In business, revenue changes might indicate temporary fluctuations or significant problems ahead. Healthcare professionals face similar challenges-without proper information gathering, diagnosis becomes impossible.
The COVID-19 pandemic spread rapidly because leaders failed to see and understand the virus properly. The critical mistake was assuming people were only contagious when symptomatic, as reflected in CDC guidance from March 2020. This fatal error occurred because past experience with SARS in 2003 taught the wrong lesson-that screening for symptoms like fever was sufficient. By missing asymptomatic transmission, all subsequent measures failed to stop the virus from propagating.
While some might consider this a "Black Swan" event, the signs of such disasters are actually lurking below the surface of everyday successes. We can prepare by looking deeper and not waiting for catastrophe to strike.
Глава 3
The Sequence of Reliability: A Revolutionary Framework
After decades of studying catastrophes across industries - from nuclear power plants to healthcare systems to aviation disasters - Griffith discovered a fundamental pattern that transcends specific fields: a precise sequence that enables consistent reliability over time. This "Sequence of Reliability" provides both practical and theoretical value for reducing risk and achieving better results in any context, whether in business operations, healthcare delivery, or personal development.
The sequence has two key steps that must be followed in order:
First, see and understand risk - this requires developing both the vision to recognize potential dangers and a deep knowledge of how these risks can manifest and harm us. This understanding must be comprehensive, considering not just obvious immediate threats but also subtle, systemic risks that can compound over time. For example, in healthcare, this means understanding both acute risks (like medication errors) and chronic systemic risks (like communication breakdowns between departments).
Second, manage reliability in this specific sequence:
1. Systems (to become effective and resilient) - establishing robust processes, protocols, and infrastructure
2. Humans (performance and behavior) - addressing training, decision-making, and human factors
3. Organizations (to achieve sustainment and become predictive) - creating cultures and structures that support reliability
This order matters profoundly and represents a paradigm shift from traditional approaches. Most organizations get it backward-focusing on human behavior before fixing flawed systems, or implementing organizational changes without understanding the underlying risks. This leads to frustration, waste, and continued failures.
Consider dieting, traditionally considered a function of willpower. Focusing on willpower alone is out of sequence and often doomed to fail. Without an effective dietary system that works for our individual metabolism, no amount of willpower will lead to sustainable weight loss, often resulting in the "yo-yo effect." For effective weight management, we must first understand the science of nutrition and metabolism, then select the right system (diet) that accounts for individual factors like schedule, preferences, and lifestyle, and only then focus on managing human performance (willpower and habits).
Similarly, with autonomous vehicles, we must first understand the complex web of psychological, physical, and emotional effects as humans adapt to automation. Drivers will face confusion between autopilot systems and manual control, similar to what commercial aviation experienced in the late twentieth century when cockpit automation led to several accidents due to mode confusion. The aviation industry learned that understanding these risks and developing appropriate systems must precede training pilots and establishing organizational protocols.
The sequence applies equally in high-reliability organizations like nuclear power plants, where system design must precede operator training and organizational policies. When managing risk, we often prematurely focus on human behaviors before properly understanding the dangers and establishing robust systems. But following the correct sequence - understanding, systems, humans, then organizations - is crucial for achieving optimal outcomes and sustainable reliability.
This framework has been validated across numerous industries, from healthcare to manufacturing, demonstrating that reliability isn't about perfect performance but rather about following the right sequence in risk management and improvement efforts.
Глава 4
System Reliability: The Foundation of Safety
Systems fail-from McDonald's ice cream machines to power grids to bridges-sometimes with fatal consequences. For leaders, understanding how systems fail and work is crucial for both safety and business success.
System performance is shaped by interconnected factors: system design and degradation, resource matching, capacity and operational load, external factors, and human performance. All systems deteriorate over time: hardware wears out, software becomes outdated, and procedures lose effectiveness. System capacity must be designed to handle operational load, as seen in tire ratings, computer disk space, and classroom sizes.
Among these factors, system design stands as the most crucial element for both effectiveness and resilience. Engineers manage system performance through three sequential strategies:
1. Barriers prevent failures by reducing threats or limiting risky human performance. They come in many forms: physical (fences), electrical (ground fault interrupters), software (passwords), biological (antibiotics), permanent (speed limits), temporary (school zone restrictions), psychological (surveillance cameras), and regulatory (laws and procedures). Well-designed barriers can dramatically lower risk without changing human behavior, as with breakaway gas pump hoses that prevent fuel leaks when drivers forget to remove the nozzle.
2. Redundancies provide reliability through parallel working components and backups. During normal operations, parallel components like dual truck wheels offer extra reliability. When things go wrong, backups like spare tires become essential. Aviation exemplifies redundancy with dual engines, tires, landing lights, windows, wipers, pilots, multiple hydraulic systems, and electrical systems. These redundancies work best when truly independent-contaminated fuel from the same source affects both tanks.
3. Recoveries provide additional protection after barriers and redundancies fail, potentially reversing harm. These include physical (parachutes, spare tires), software (system-restore functions), biological (medication antidotes), and people/processes (lifeguards, roadside assistance). While recoveries often require fewer initial resources than barriers or redundancies, they're implemented after something has already gone wrong.
A hospital story illustrates this approach: A nurse at a cardiac unit heard an alarm and systematically turned off equipment to locate the source. After finding it was a false alarm, she forgot to reactivate the cardiac monitor. When the patient later went into cardiac arrest, no alarm sounded and the patient died.
Rather than firing the nurse, the patient safety officer recognized this as both a human and system failure. The officer approached the cardiac monitor manufacturer with a simple solution: program the device to automatically reactivate after being turned off. This recovery strategy reduced the likelihood of similar incidents.
Глава 5
Human Reliability: Managing Our Inevitable Mistakes
We're all human-unique products of our biology, environment, and experiences-yet we share one universal trait: we all make mistakes without exception. Though human failures are inevitable, they can be managed by understanding how people perform, what motivates them, and what influences their behaviors.
Human reliability exists in tension with system reliability. Often, the more reliable our systems become, the less reliable humans grow within them. When people rely on automation, they drift into complacency and lose proficiency in manual skills-from remembering phone numbers to flying aircraft without autopilot. The goal isn't perfection but resilience-maintaining human proficiency even when automation handles normal operations, so people can recover when systems fail.
Human reliability depends on understanding interconnected factors that shape performance before addressing behaviors. These factors include:
• Knowledge, Skills, Abilities, and Proficiency (KSAPs) that evolve throughout our lives
• System Factors including training, policies, equipment, and environmental conditions
• Personal Factors encompassing health, personal conflicts, and past experiences
• Culture that powerfully shapes behavior through organizational values, peer pressure, and leadership
Daniel Kahneman identifies two fundamental thinking modes that affect our reliability:
• System 1 (fast, automatic): We operate on autopilot with relaxed facial muscles, making quick decisions based on training and instinct. While efficient, System 1 is prone to errors.
• System 2 (slow, deliberate): Involves concentrated, methodical thinking, often reflected in scrunched facial expressions. Though more logical, System 2 requires greater mental effort, so we use it less frequently.
Human errors come in distinct forms:
• Slips: Physical or cognitive actions we didn't intend to make-like hitting "send" on an email prematurely
• Lapses: Inadvertent failures to meet expectations-such as forgetting to turn off the stove
• Mistakes: Errors where the action was intended but the result was not-like an umpire making an incorrect call
At-risk choices are different from errors-they're intentional behaviors that increase risk where that risk isn't recognized or is mistakenly believed justified. Examples include speeding, eating while driving, or handling hot food without protection. The challenge in managing at-risk choices is that they usually produce positive results, reinforcing the behavior until tragedy strikes.
Steve Irwin's case perfectly illustrates this. His decision to bring his one-month-old son into a crocodile pit sparked public outrage, but Irwin was genuinely perplexed by the reaction. As he explained to Larry King, he never imagined any outcome other than success-he knew the crocodile's temperament, had prepared extensively, and had done similar things with his daughter for years. His tragic death in 2006 from a stingray attack demonstrates how repetitive at-risk choices eventually produce different outcomes as conditions change.
Глава 6
Organizational Reliability: Creating Sustainable Success
Organizations are socio-technical combinations of people working within systems. Their reliability depends not just on individual behaviors but on how these elements interact within structured systems. Like a professional baseball team with its farm system, managers, coaches, and players, organizations operate through interconnected subsystems where human performance contributes to success or failure.
Organizations face competing priorities driven by multiple values-principles that guide behavior and judgment. While businesses might list core values like customer service, safety, privacy, and cost control, no single value consistently dominates in all circumstances. Competing priorities are inevitable-employees develop workarounds as circumstances change, even while core values remain intact.
Leaders face two fundamental challenges: they can lead in the wrong direction as easily as the right one, and when leading correctly, they must understand how systems and people achieve reliability. Effective leadership requires coding reliability into organizational DNA by following the Sequence: seeing and understanding risk, then managing system and human performance before addressing organizational performance.
Culture is a dynamic, ever-evolving force affecting organizational performance at every level. Like Uber's cautionary tale demonstrates, toxic culture can undermine even the most innovative business model. Organizations contain multiple micro-cultures coexisting with macro-cultures-varying by department, work unit, and even shift. Rather than seeing this diversity as chaos, successful organizations recognize that cultural diversity provides a system design advantage similar to "hybrid vigor" in ecological systems.
Biases distort our perception and lead to dangerous conclusions:
• When we punish at-risk choices only after bad outcomes occur, we fail to address the countless similar behaviors happening without incident
• The most dangerous form of outcome bias is underreacting when nothing bad happens-the "no harm, no foul" mentality
• Professional bias creates double standards across hierarchies
• Fundamental attribution error causes us to judge others more harshly than ourselves
• Defensive attribution hypothesis makes us attribute more blame to people different from us
• Rule biases create tension between managers who overvalue compliance and employees who undervalue risks
• Normalization of deviance allows dangerous practices to become accepted
• Confirmation bias leads us to perceive what we expect
NASA's space shuttle program offers critical lessons in organizational reliability. Despite completing 133 successful missions, the Challenger and Columbia disasters revealed how organizational culture contributed as much as technical factors to these failures. Remarkably, NASA had pioneered socio-technical probabilistic risk assessments and had even quantified many specific risks. Yet competing priorities of supporting the International Space Station while being fiscally responsible created organizational pressure that contributed to the disasters.
Глава 7
Predictive Reliability: Looking Beyond Past Events
Predictive reliability means looking beyond past events to anticipate future risks. Just as we wouldn't use last month's weather patterns to predict next year's climate, we can't rely solely on past performance to anticipate future risks.
Organizations identify risk through progressively more valuable strategies:
1. Accident investigations
2. Audits and inspections
3. External reports
4. Employee reporting systems
5. Informal observations
6. Digital surveillance systems
7. Predictive risk modeling
Each method has limitations and is subject to interpretation bias. The goal is developing a complete risk picture by combining complementary strategies while recognizing their inherent limitations.
Early organizational improvement models described accidents as "links in a chain," suggesting we could prevent failures by removing a single link. This dangerous oversimplification offers no rational basis for determining which links to remove first. James Reason's Swiss cheese model added a third dimension, viewing accidents as the result of latent failures and hazards aligning like holes in Swiss cheese. But this still leaves us questioning which holes to plug first.
Predictive risk modeling emerged in the 1970s from nuclear safety concerns, using fault trees to map system failures mathematically. These trees illustrate multiple pathways to failure with probability estimates for each branch. The approach starts by narrowing broad risks to specific subcategories. Fault trees use AND gates (requiring all conditions to occur) and OR gates (requiring any condition to occur) to calculate failure probabilities. This creates a quantifiable risk picture highlighting the most likely pathways to accidents-not just the one that happened yesterday.
The Metrolink train crash analysis reveals how we must examine both human and system factors. The contributory map shows the crash resulted from interconnected factors: the engineer's texting (behavioral choice) led to running a stop signal (human error), which caused the collision because trains were traveling in opposite directions on the same track. While management might focus on preventing texting through monitoring or punishment, a predictive approach would identify multiple ways signals could be missed and implement system solutions like positive train control that would prevent collisions regardless of human error.
Глава 8
Flipping the Iceberg: From Reactive to Proactive
After witnessing Flight 191's crash, Griffith worked with scientists developing lidar wind shear detection systems-technology that could measure invisible wind patterns ahead of aircraft, allowing pilots to take preventive action. Their airborne lidar system became the first truly predictive wind shear detection tool, later tested on the space shuttle and now used worldwide.
Building on this experience, Griffith developed the Aviation Safety Action Program (ASAP), which transformed aviation safety by encouraging frontline workers to report risks without fear. Critics called it a "get-out-of-jail-free card," but couldn't argue with the program's success. Following industry-wide adoption, a clearer picture of risks below the waterline emerged, with the industry identifying precursors like pilots falling asleep on long flights, automation confusion, and risky choices driven by operational pressures.
The US aviation industry's remarkable safety improvements over decades have advanced reliability science with lessons applicable to other industries. In 1995, following public concern over fatal commercial airplane accidents, Transportation Secretary Federico Pena declared, "The industry must make it clear to the public we will not settle for anything less than zero accidents." Two years later, the White House Commission on Aviation Safety and Security revised the goal to an 80% reduction in fatal accidents over ten years-by 2005, they achieved 78%.
This success came through industry-wide collaboration between airlines, labor, regulators, and manufacturers. The Commercial Aviation Safety Team established consensus-based, data-intensive reviews of worldwide accidents, proposing joint solutions across stakeholders.
To address limitations in translating these lessons to other industries, Griffith developed the Collaborative Just Culture (CJC) program-an evidence-producing, standardized approach combining ASAP principles with the Sequence of Reliability. CJC requires documented elements including executive leadership commitment, employee involvement, written policies, training requirements, fact-gathering tools, and a sustainment plan. Most importantly, it implements a Triad review process with representatives from management, HR, and safety/quality/risk to ensure balanced organizational responses through unanimous consensus.
Building on this foundation, Griffith created Collaborative High Reliability (CHR)-an evidence-based approach drawing from aviation safety programs, Collaborative Just Culture, integrated engineering and behavioral sciences, and ISO 9001 quality principles. CHR uses a clear taxonomy distinguishing between activities, processes, programs, systems, and integrated systems. This framework allows organizations to properly document, align, and measure reliability efforts.
To ensure proper implementation, Griffith partnered with DNV (Det Norske Veritas), a leading global certification body, to create the world's first independently audited high-reliability model. The DNV CHR audit process is both independent (third-party verification) and proficiency-based (requiring demonstrated competency over time).
Глава 9
The Path Forward: Applying Reliability Science
The challenges we face today-from climate instability to pandemic prevention to employee burnout-require applying the Sequence of Reliability. Organizations must look below the surface using predictive methods before disasters occur.
When assessing risk and considering solutions, we should follow three steps:
1. Determine the probability that system design and human response strategies will effectively mitigate risk
2. Assess operational effects to predict intended and unintended consequences
3. If deemed effective, provide operational and financial support for improvements
The pattern is clear-look first to strong system-focused solutions rather than concentrating on human performance alone. Reliable organizations invest in sustainable success by managing risk proactively rather than reactively, becoming good at preventing what they never intend to do.
A gas and electric utility company with over 800 large vehicles faced a problem with drivers backing into things. Despite installing backup cameras, mirrors, proximity sensors, and using spotters, accidents continued. The assessment team discovered the drivers were professionals in their gas and electric work, not driving professionals. With twelve competing sensory inputs in most trucks, drivers couldn't process all information simultaneously. The solution was to remove selected sensory inputs, adopt a "Stop, Scan, then Primary" method of backing up, and reduce multitasking while in motion. After implementing these recommendations, reported accidents dropped to zero in the first year and just two in the second year.
As Richard Feynman said, "Our responsibility is to do what we can, learn what we can, improve the solutions, and pass them on." The science is no longer hidden, the sequence matters, and positive results are sustainable. By applying the Sequence of Reliability, organizations can achieve breakthrough performance and navigate an increasingly complex world with confidence.
Глава 10
Transforming Risk Management for a Safer Future
The key principles for managing risk begin with developing the ability to see and understand what isn't immediately obvious, then systematically improving systems, effectively managing people and organizations, and sustaining success through unwavering commitment. Experience alone proves insufficient for avoiding risks you haven't directly encountered-you need sophisticated tools and frameworks to identify potential catastrophes before they materialize. This requires moving beyond traditional risk matrices to employ advanced predictive modeling and scenario planning techniques.
Systems typically fail due to multiple interconnected factors, with system design being the most critical element. Complex systems often harbor hidden vulnerabilities that only become apparent after failure occurs. Building in multiple layers of protection through barriers, redundancies, and recovery mechanisms significantly improves reliability. These might include physical barriers, procedural controls, and automated safety systems working in concert. For example, in aviation, multiple backup systems exist for critical functions, and procedures are designed with numerous cross-checks and verification steps.
The greatest risks emerge not from simple human errors but from the complex choices people make under various influences including time pressure, resource constraints, and competing priorities. Organizations must systematically identify these competing priorities and cognitive biases that can compromise performance. This knowledge should inform the design of resilient systems that account for human limitations while recognizing that risk factors evolve over time. For instance, production pressure often competes with safety protocols, requiring careful balance and clear decision-making frameworks.
Historical accidents provide important but limited guidance for future prevention, as the nature of risks constantly evolves with technological advancement and organizational complexity. This necessitates implementing multiple complementary strategies and fostering genuine stakeholder collaboration across all levels. Modern challenges require rigorous application of the Sequence of Reliability, which means developing the capability to see below the waterline of visible events to identify root causes and implementing Collaborative High Reliability practices with independent validation from external experts.
The Sequence of Reliability offers a comprehensive path forward not just for high-consequence industries like healthcare, aviation, and energy, but for any organization seeking sustainable success in an increasingly uncertain world. By understanding the hidden patterns of risk and applying this scientific approach, leaders can transform their organizations from reactive to predictive modes of operation. This transformation requires sustained commitment to building robust systems, developing people's capabilities, and fostering a culture of openness and learning. Successfully implementing these principles ultimately saves lives, protects fortunes, and preserves reputations by preventing catastrophic failures before they occur.