1장
When Numbers Speak: The Hidden Patterns That Shape Our World
Have you ever wondered why traffic jams form even when there's no accident ahead? Or why Disney World's wait times feel shorter than they actually are? In his groundbreaking book "Numbers Rule Your World," Kaiser Fung pulls back the curtain on how statistics silently shape our everyday experiences. Unlike many books that focus on statistical misuse, Fung takes the refreshing approach of examining what happens when numbers tell the truth-and why we still struggle to listen.
Since its publication, the book has become required reading in university statistics programs worldwide and has been praised by Nobel laureate Daniel Kahneman as "a rare gem that makes statistical thinking accessible and engaging." Through real-world examples ranging from theme parks to credit scores to disease outbreaks, Fung demonstrates how statistical thinking solves practical problems in ways that transform industries and save lives.
2장
The Unpredictable Everyday: Why Averages Aren't Enough
What frustrates commuters isn't average travel time but its maddening unpredictability. Julie Cross faces "a daily gamble" between a reliable 10-minute interstate trip versus a potentially faster but inconsistent Cedar Avenue route that can vary from 8 to 25 minutes. This daily decision mirrors countless others made by commuters worldwide, where the uncertainty of travel time, not its duration, causes the most stress. Similarly, Disney World visitors spend 3-4 hours in queues during an 8-9 hour visit, with wait times swinging from 15 minutes to over 2 hours for popular attractions like Space Mountain, depending on time of day, weather, and seasonal factors.
This insight-that variation matters more than averages-represents a fundamental principle of statistical thinking. Belgian statistician Adolphe Quetelet invented the concept of "l'homme moyen" (the average man) in 1831, revolutionizing social sciences but creating a fundamental problem: averaging erases diversity. For instance, designing a cockpit for the "average pilot" in the 1950s proved disastrous - not a single pilot among 4,063 measured fit within the average of all 10 physical dimensions. This demonstrated how averages can create solutions that work for nobody.
Minnesota's Department of Transportation tackled highway variability with "ramp metering"-stop-go lights on highway entrances that regulate traffic flow. The system uses sophisticated algorithms to monitor real-time traffic density, speed, and flow rates. The results were impressive: Seattle saw traffic volume increase 74% while average journey time was halved during peak hours. Portland achieved similar success, reducing travel times by 20.7% after implementing metering. This "freeway congestion paradox" works because once congestion starts, both speed and carrying capacity plummet dramatically-from 2,000 to 1,200 vehicles per hour. Ramp metering maintains optimal highway speeds by throttling vehicle influx, much like controlling water flow through a pipe.
Meanwhile, Disney's "Imagineers" discovered something more powerful than reducing actual waiting time: managing perception of waiting. Research shows perceived waiting differs from actual waiting-mirrors in elevator lobbies distort time perception because people don't count time spent looking at themselves as waiting. Disney masterfully shapes perceptions through multiple strategies: themed waiting areas that tell stories, interactive elements like the Haunted Mansion's ghostly library, roaming entertainers, and deliberately overestimating posted wait times by 20%. Their crowning achievement is FastPass, which transforms standing-in-line time into free time elsewhere in the park without actually reducing total waiting time. This system effectively redistributes crowd flow while making waiting feel more palatable.
Despite Mn/DOT's nationally recognized success, public perception turned against ramp metering. When the legislature mandated a six-week "meters shutoff" experiment in 2000, travel times rose 22% and crashes during merges jumped 26%. The number of rear-end collisions increased by 15%, and rush-hour traffic speeds dropped by 7 mph. Yet subjectively, drivers still preferred moving slowly on freeways to standing still at ramps, highlighting the gap between objective improvements and subjective experience. Engineers finally acknowledged this psychological dimension, implementing compromises like four-minute maximum waits and more transparent communication about waiting algorithms. The lesson? Technical solutions must address not just objective measures but also how people perceive and experience systems, as human psychology often trumps statistical efficiency.
3장
When Correlation Trumps Causation: The Power of Patterns
Credit scoring models have revolutionized the lending industry, transforming what was once a weeks-long subjective process into automated decisions delivered in seconds. These sophisticated models analyze hundreds of variables across millions of consumer records. Critics frequently demand clear causal relationships between specific behaviors and creditworthiness, but in the realm of consumer behavior, correlation proves both sufficient and necessary for effective modeling. Unlike the physical sciences with their immutable laws, human behavior is inherently complex - people are "moody, petulant, haphazard, and adaptive," making strict causation elusive.
Modern credit scoring systems process vast amounts of data - billions of data points monthly across major credit bureaus. Their resilience comes from built-in redundancy, evaluating clusters of related characteristics together. For instance, stability indicators include residence duration, job tenure, and credit history length, while financial responsibility measures encompass payment history, credit utilization, and types of credit used. This redundancy ensures that individual data errors or anomalies don't significantly impact overall scores.
The models' effectiveness stems from sophisticated pattern recognition rather than simple linear relationships. For example, someone with a short credit history but stable employment and residence might receive a better score than someone with longer credit history but frequent moves and job changes. These nuanced interpretations have withstood decades of real-world testing across economic cycles and demographic shifts.
Contrary to common perception, data inaccuracies don't uniformly harm consumers. While some receive artificially lower scores due to errors, others benefit from undeservedly higher scores - though these beneficial mistakes rarely generate complaints. The Fair and Accurate Credit Transactions Act of 2003, while well-intentioned, created opportunities for exploitation. Credit repair schemes emerged, including "piggybacking" services where individuals pay to be added as authorized users on strangers' well-maintained credit accounts, artificially inflating their scores.
Credit scoring has actually democratized lending, expanding access to previously excluded groups. Between 1989 and 2004, households earning under $30,000 annually increased their borrowing by 247%. While certain demographic groups show lower average scores, this reflects broader economic disparities rather than algorithmic bias. The complexity of modern scoring models often approves lower-income applicants who demonstrate positive patterns in other areas - applications that simpler, more rigid systems would have rejected outright.
The contrast with epidemiology is instructive. During the 2006 spinach E. coli outbreak, epidemiologists employed case-control studies comparing consumption patterns between affected and healthy individuals. They methodically traced the contamination to specific farms in California's Salinas Valley, demonstrating how causation-focused approaches serve different purposes than correlation-based models. This precision allowed targeted recalls while protecting the broader produce industry.
Statistical modeling has transformed both fields, though their approaches differ fundamentally. Epidemiology requires causation because biological mechanisms ultimately drive disease transmission. Credit scoring embraces correlation because human financial behavior emerges from countless interacting factors that defy simple causal explanation. As statistician George Box famously noted, "All models are wrong but some are useful" - the goal isn't perfect prediction but meaningful improvement over previous methods. Modern credit scoring achieves this through sophisticated pattern recognition while acknowledging inherent limitations.
4장
The Group Dilemma: When Fairness Means Treating People Differently
J. Patrick Rooney, CEO of Golden Rule Insurance and creator of health savings accounts, became an unlikely civil rights champion in the mid-1970s when he noticed his company had no Black insurance agents in Chicago. He sued the Educational Testing Service (ETS), claiming their licensing exam unfairly disqualified Black applicants. The resulting "Golden Rule settlement" pioneered techniques for identifying potentially biased test questions.
This case highlights a fundamental statistical dilemma: when should we treat groups together, and when separately? Test developers learned they couldn't directly compare students across racial groups without accounting for educational opportunity differences. Using differential item functioning (DIF) analysis, they identify questions that function differently across demographic groups with similar abilities. When a question shows a 15% performance gap between racial groups that cannot be explained by overall ability differences, it's removed from the test.
The insurance industry faces similar grouping challenges. Bill Poe's Florida insurance company collapsed after the unprecedented hurricane seasons of 2004-2005, despite his forty years of experience. The eight hurricanes that struck Florida caused $36 billion in losses, erasing all industry profits since 1993. This catastrophic outcome revealed fundamental flaws in hurricane insurance models.
Unlike auto insurance, where only a small portion of policyholders file claims in any year, natural disaster insurance faces concentrated geographic risk. When hurricanes strike, they affect massive portions of the customer base simultaneously. Poe's fatal mistake was concentrating risk in South Florida, resulting in nearly 40% of customers filing claims at once.
The industry had also misinterpreted the concept of "100-year storms," failing to recognize that such events have a cumulative probability over time. Statistics show that over ten years, the chance of experiencing a 100-year hurricane is roughly 10%, not 1%, explaining why powerful hurricanes could strike Florida in succession.
As the market adjusted, insurers began stratifying risk pools, separating inland and coastal properties. National companies created Florida-only subsidiaries to protect corporate parents, while reinsurers doubled rates. The state-run Citizens Property Insurance Corporation inherited Poe's 330,000 customers while already running a $1.7 billion deficit. To cover these losses, Florida levied numerous "assessment fees" on all insurance policies statewide.
This created a perverse situation where inland residents continued subsidizing coastal properties, but now through government-mandated fees rather than private insurance pools. The fundamental question in both testing and insurance is whether groups should be lumped together or treated separately-the dilemma of being together.
5장
The Asymmetry of Error: When All Mistakes Aren't Created Equal
In the early 2000s, Major League Baseball finally implemented steroid testing amid growing public concern about performance-enhancing drugs. Unlike the World Anti-Doping Agency's stringent standards, MLB's testing program developed incrementally, with players like Mike Lowell demanding "100 percent accuracy" to prevent false positives-clean athletes wrongly accused of doping. With average MLB salaries reaching $2.5 million in 2005, players feared career-destroying false accusations more than the possibility of cheaters escaping detection.
Meanwhile, in Iraq and Afghanistan, military interrogators faced a different challenge: screening local job applicants for insurgent connections. When portable lie detectors arrived in 2007, they promised to remove human judgment from counterintelligence work. With American lives at stake, the military calibrated these devices to minimize false negatives-insurgents passing screening-potentially at the cost of increased false positives.
Both steroid testing and lie detection face an unavoidable statistical trade-off between false positives and false negatives. As with a baseball hitter who can swing more aggressively (risking strikeouts) or less aggressively (hitting fewer home runs), detection systems can be calibrated but cannot simultaneously minimize both error types. When one error type is more visible or costly, systems become asymmetrically focused, often neglecting the less apparent error type.
The polygraph machine-a collection of medical tools measuring breathing, blood pressure, pulse rate, and skin conductivity-detects anxiety, not deception itself. Created by William Marston (who later created Wonder Woman with her truth-compelling Magic Lasso), the polygraph requires skilled interpretation by examiners, typically retired law enforcement personnel with specialized training.
Despite widespread public acceptance, polygraphs remain inadmissible in most U.S. courts, failing the "general acceptance" standard established in the 1920s. Scientific reviews consistently warn about their unreliability, especially for screening purposes. Congress has sent mixed signals-banning polygraph screening in private employment while allowing government agencies free rein.
The Angela Correa murder case demonstrates the polygraph's power as a confession tool. When 15-year-old Angela was found raped and murdered in 1989, police quickly focused on 16-year-old Jeffrey Deskovic despite DNA evidence specifically excluding him. After being told he failed a polygraph test, Deskovic confessed. This confession secured his conviction despite contradictory scientific evidence. He served sixteen years before being exonerated.
When the U.S. Army deployed the Preliminary Credibility Assessment Screening System (PCASS) in Iraq and Afghanistan in 2007, statistician Stephen Fienberg was appalled. As technical director of the National Academy of Sciences' 2002 report rejecting polygraph technology as unreliable, particularly for security screening, he saw this $2.5 million investment as a blatant disregard for scientific consensus.
The Army's PCASS calibration reveals their overwhelming fear of false negatives-they'd rather wrongly flag innocent people than miss a single insurgent. With the system set to pass fewer than 50% of subjects, the statistical consequences are severe: for every true insurgent identified, approximately 93 innocent people are falsely classified as deceptive. This asymmetric approach stems from the belief that even one undetected insurgent could be devastating.
6장
When Lightning Strikes: Making Sense of Rare Events
We approach extremely rare events with contradictory attitudes-eagerly playing lotteries while fearing air travel, despite similarly minuscule odds for both. This cognitive dissonance reveals how emotion, rather than probability, often drives our risk assessment. For instance, the annual risk of dying in a plane crash is roughly 1 in 11 million, while the odds of winning a major lottery are typically 1 in 14 million, yet we perceive these probabilities very differently.
When multiple plane crashes occurred in the "Corridor of Conspiracy" near Nantucket in the late 1990s, many rejected randomness as an explanation, believing some hidden cause must exist. This statistical reasoning isn't irrational-we naturally seek patterns and causes, especially in tragic events. However, aviation expert Arnold Barnett demonstrates that when viewed against millions of safe flights-approximately 100,000 commercial flights operate daily in the US alone-these rare crashes appear truly random. Barnett's analysis showed that clustering of accidents in time or space is mathematically expected, just as a truly random coin flip sequence will naturally contain streaks.
Meanwhile, statistician Jeffrey Rosenthal used similar statistical testing to uncover genuine lottery fraud in Ontario, where store insiders won 200 major prizes when probability suggested they should have won only 57. This dramatic disparity-a deviation of more than 140 unexpected wins-signaled systematic abuse. The chapter details how Ontario store owner Phyllis LaPlante defrauded 82-year-old Bob Edmonds of his $250,000 lottery win by convincing him he'd only won a free ticket when his numbers came up. When LaPlante claimed the prize herself, it triggered an "insider win" investigation, but she initially passed by providing old tickets with Edmonds' regular numbers. This case highlighted how vulnerable elderly players were to fraud and led to major reforms in lottery ticket validation procedures.
Rosenthal's lottery fraud investigation spread across Canada, uncovering suspicious patterns in multiple provinces. In British Columbia, officials dismissed concerns despite evidence showing insider win rates three times higher than expected. In New Brunswick, an audit revealed store owners won 37 out of 1,293 major prizes when statistically expected to win fewer than 4-a deviation so extreme it couldn't be explained by chance. The Western provinces showed similar patterns-insiders won twice as many major prizes as expected, with odds of 1 in 2.3 million that this occurred by chance. These findings led to widespread reforms, including mandatory video surveillance of ticket validations and automated prize claim tracking systems.
Statisticians evaluate patterns against complete backgrounds, not just anomalies, using sophisticated tools like regression analysis and probability theory. Their worldview holds that "rare is impossible"-they reject explanations requiring extremely improbable events, following Occam's Razor principle that simpler explanations are usually correct. When analyzing the Minnesota ramp metering experiment, statisticians concluded the program was effective because the alternative explanation-that a rare event caused the 22% increase in travel time when meters were shut off-was statistically improbable. This approach to rare events, combining rigorous mathematical analysis with practical skepticism, helps separate genuine patterns from random fluctuations in fields ranging from public safety to fraud detection.
7장
Statistical Thinking: A Different Way of Seeing the World
Psychologist Daniel Kahneman emphasizes that statistical thinking demands deliberate mental effort because our brains have evolved for quick, intuitive judgments rather than probabilistic reasoning. Throughout the book, we've explored five fundamental aspects of this counterintuitive way of thinking:
First, the discontent of being averaged-look beyond averages to understand variability. Bernie Madoff's investors learned this lesson catastrophically when their supposedly stable 12% annual returns proved to be an elaborate fraud. Business metrics like compound annual growth rates (CAGR) can be similarly deceptive, masking dramatic year-to-year swings that significantly impact operations and cash flow. For instance, a company showing 10% CAGR might experience years of 25% growth followed by 5% declines. Effective solutions like Disney's FastPass system and highway ramp meters succeed precisely because they address underlying variability rather than just managing average wait times or traffic flow.
Second, the virtue of being wrong-choose useful over true. As statistician George Box famously declared, "All models are false but some are useful." Great statisticians embrace this inherent fallibility, seeking models that best fit available evidence without claiming absolute truth. In credit scoring, correlation-based models have proven remarkably effective despite not explaining causation. FICO scores, for instance, successfully predict creditworthiness using factors like payment history and credit utilization, even though they don't explain why these correlations exist.
Third, the dilemma of being together-compare like with like. Following the devastating 2004-2005 hurricane seasons, Florida insurance companies recognized that coastal and inland properties had become fundamentally different risk categories. Pooling them together would unfairly burden inland residents with coastal property risks. Simpson's paradox further illustrates this complexity: when comparing test scores between demographic groups, combining different ability levels can create apparent gaps even when none exist within each level. For example, UC Berkeley's graduate admissions appeared biased against women overall, but examination by department showed either no bias or bias in women's favor.
Fourth, the sway of being asymmetric-balance two types of errors. Statistical errors manifest as false positives (false alarms) and false negatives (missed opportunities). Improving detection of one inevitably increases the other. Medical screening programs often face this trade-off: more sensitive cancer screening catches more cases but also increases false alarms, while stricter criteria reduce false positives but miss more actual cases. Decision-makers frequently focus on avoiding whichever error brings negative publicity, overlooking the hidden costs of the alternative.
Finally, the power of being impossible-reject what's too rare to be true. When analyzing global airline safety, statistician Arnold Barnett couldn't find sufficient evidence to reject the hypothesis that developed and developing-world carriers were equally safe, despite apparent differences. In contrast, mathematician Jeffrey Rosenthal's analysis of Ontario's Encore lottery winners revealed that store owners' winning frequency was so improbable (one in a quindecillion assuming fair play) that systematic fraud was the only plausible explanation. This principle of rejecting the virtually impossible helps detect everything from scientific fraud to financial manipulation.
8장
The Applied Scientist: Where Theory Meets Reality
Throughout the book, we've met unsung heroes who excel not through invention but through adaptation, refinement, and perseverance in applied science. These applied scientists differ from pure scientists in several ways. They make decisions affecting real lives with practical goals-societal, psychological, or financial. They work under tight time constraints rather than pursuing perfection indefinitely. They adapt elegant theoretical models by accounting for messy real-world details.
Len Testa's Disney theme park touring plans demonstrate the power of applied statistics. His team collects wait times at every ride every thirty minutes throughout the year, walking eighteen miles daily through the parks. Their algorithm identifies that ride popularity and time of day matter most (rated 10), followed by crowd level (9), holiday (8), early-entry morning (5), day of week (2), and weather (1). Without explaining why certain times are busier, the model successfully helps visitors experience 70% more attractions and save 3.5 hours of waiting time.
Arnold Barnett's airline safety analysis demonstrates the importance of proper group comparison. While developing-world carriers accounted for 74% of crash fatalities despite operating only 18% of worldwide flights, Barnett discovered this statistic was misleading. Since American travelers only choose between developed and developing-world carriers on "between-worlds" routes, he isolated these relevant routes for comparison. There, developing-world carriers suffered 55% of fatalities while making 62% of flights-indicating they were actually safer than developed-world airlines.
Banks' credit scoring systems perfectly illustrate asymmetric error costs. Loan officers prioritize avoiding false negatives (bad loans) over false positives (rejected good applicants) because defaults directly impact the bottom line while missed sales opportunities remain invisible. During the early 2000s credit boom, low interest rates and economic expansion changed this calculus-the opportunity cost of missed sales increased while default risks seemed lower. Banks relaxed lending standards, inevitably increasing false negatives. By the late 2000s, many institutions collapsed under the weight of these delinquent loans.
Successful applied scientists understand decision-making contexts and can translate logical arguments for intuitive thinkers. They recognize that technical solutions must address not just objective measures but also how people perceive and experience systems. The Minnesota ramp meter experiment showed that engineers valued reduced travel time while commuters disliked waiting at ramps more than stop-and-go traffic. Similarly, Disney discovered that managing perception of waiting time was more powerful than reducing actual waiting time.
Numbers already rule your world-understanding how applied scientists use statistical thinking can help you make better everyday decisions. By focusing on variability rather than averages, embracing useful models over perfect ones, comparing like with like, balancing different types of errors, and rejecting the impossibly rare, you can navigate an increasingly complex world with greater clarity and confidence.