第1章
When Judgment Goes Awry: The Hidden Epidemic of Noise
Imagine you've been arrested for a minor offense. The judge hearing your case had a bad night's sleep and just watched his favorite football team lose. Without realizing it, he sentences you to three years in prison-while another defendant with an identical case received only six months from a different judge the day before. This isn't a dystopian fiction; it's the reality of what Nobel Prize-winning psychologist Daniel Kahneman calls "noise" in our judgment systems. Kahneman's book "Noise," co-authored with Olivier Sibony and Cass Sunstein, has become a sensation among business leaders, policymakers, and anyone making consequential decisions. Endorsed by Bill Gates as "one of the most important books I've ever read," it exposes a hidden epidemic of inconsistency that costs organizations billions while creating profound unfairness. The book's insights have transformed how companies like Google conduct interviews and how hospitals diagnose patients. Beyond its practical applications, "Noise" challenges our fundamental understanding of human judgment, revealing that where we expect consistency, we often find a troubling lottery.
第2章
The Two Faces of Error: Bias vs. Noise
When we think about errors in judgment, we typically focus on bias-systematic deviation in a particular direction. But there's another type of error that's equally damaging yet largely invisible: noise, or random scatter in judgments that should be identical. Using a shooting range metaphor, bias is like consistently missing the target in the same direction, while noise is like shots scattered widely around the bull's-eye. The critical difference is that noise can be measured without even knowing where the target is-simply by observing how widely dispersed the shots are.
This distinction isn't merely academic. Noise manifests alarmingly across numerous domains: doctors disagree on diagnoses when examining the same patient; asylum decisions depend heavily on which judge hears the case; fingerprint examiners reach different conclusions about the same prints; and insurance underwriters quote wildly different premiums for identical risks.
Consider a landmark study of criminal sentencing that revealed shocking disparities: two first-time offenders convicted of similar check fraud received sentences of fifteen years versus thirty days; similar embezzlement cases resulted in 117 days versus twenty years imprisonment. Judge Marvin Frankel condemned these "arbitrary cruelties" as unacceptable in a government of laws. His work catalyzed formal studies confirming "astounding" variations in sentencing, with factors as irrelevant as temperature, judges' hunger levels, local football team losses, and even defendants' birthdays significantly influencing outcomes.
While the Sentencing Reform Act of 1984 attempted to address this problem through mandatory guidelines, the Supreme Court later made them merely advisory, essentially returning to what Frankel called "law without order." This pattern-recognizing noise, attempting to address it, then backsliding-repeats across institutions, partly because we find it difficult to acknowledge that professionals we trust make wildly inconsistent judgments.
What makes noise particularly pernicious is that unlike bias, which can sometimes be detected by looking for patterns, noise remains largely invisible without deliberate measurement. Most organizations and professionals have no idea how much noise affects their judgments-and when they find out, they're typically shocked.
第3章
The Noise Audit: Revealing the Invisible Problem
The authors' first encounter with organizational noise occurred at an insurance company where executives doubted noise could significantly impact their business. A simple "noise audit" proved them wrong. The company employed underwriters who set premiums and claims adjusters who estimated settlement costs-both making consequential financial judgments.
To measure variability, multiple employees independently evaluated identical case files. Before seeing the results, executives guessed differences would be around 10%-a tolerable level of variation. The actual results shocked them: for underwriters, the median difference between judgments was 55%-meaning when one underwriter might quote $9,500, another would quote $16,700 for identical risk. Claims adjusters showed 43% median difference. This "system noise" represented hundreds of millions in annual losses through overpriced quotes losing business and underpriced ones creating losses.
This pattern repeats across industries. In an asset management firm's noise audit, 42 experienced investors showed 41% median noise when valuing the same stock. The magnitude of noise consistently exceeds what professionals expect, partly because organizations naturally minimize exposure to disagreements. Many actively avoid conflict-like the school that stopped showing first readers' ratings of applicants because it "resulted in so many disagreements."
Unlike diversity in matters of taste or preference, this variability is entirely unwanted. Customers expect consistent judgments from organizations, not to be unknowingly entered into a lottery where their premium depends on which employee happens to handle their case. System noise matters because errors don't cancel out; they add up. If one policy is overpriced and another underpriced, the system has failed twice.
The noise audit methodology provides a powerful diagnostic tool. By having multiple professionals independently judge the same cases, organizations can measure system noise without needing to know the "correct" answer. This approach reveals not just how much noise exists but also its components-level noise (some judges being consistently harsher than others) and pattern noise (judges ranking cases differently).
The reality is stark: wherever judgment exists, noise exists too-and in greater amounts than we realize.
第4章
Your Mind Is a Measuring Instrument
When making judgments, we're essentially using our minds as measuring instruments. Just as a thermometer assigns temperature values to objects, our judgments assign values on various scales-whether sentencing criminals, pricing insurance, or diagnosing illnesses. Like all measurements, judgments contain error-some bias, some noise-and perfect accuracy remains unattainable.
Professional judgments occupy a middle ground between factual questions (where reasonable people should agree perfectly) and matters of taste (where disagreement is acceptable). They're characterized by an "expectation of bounded disagreement"-reasonable people might disagree, but not wildly so. The acceptable range depends on the problem's difficulty, with agreement typically easier on absurd judgments than on reasonable ones.
Making a judgment involves several cognitive processes that introduce variability. Consider evaluating a CEO candidate: we selectively attend to certain details while ignoring others-some focus on his Harvard background, others on his abrasive style. We then informally integrate these cues into an overall impression without following a structured process. Finally, we convert this impression into a specific number on a scale, often through an intuitive process where numbers "feel right" rather than through explicit calculation. Each step introduces noise into our judgments, explaining why MBA students' estimates of the same candidate's success probability ranged from 10% to 95%.
When making judgments, we're not actually trying to match an external outcome (which often doesn't exist yet), but rather seeking an internal signal of coherence-that feeling when our answer feels right because it comfortably fits with the evidence. This "internal signal of judgment completion" occurs whether the judgment is verifiable or not.
Noise in judgments indicates problems in different ways. In predictive judgments (like medical diagnoses), disagreement means at least one person must be wrong, potentially causing serious consequences. In evaluative judgments (like sentencing), noise violates expectations of fairness and consistency-turning justice into a lottery. Even when unfairness isn't the primary concern, system noise represents inconsistency that damages credibility, as when similar customer complaints receive dramatically different responses.
The good news is that we can measure noise without knowing the "true" value of a judgment-we simply need multiple judgments of the same problem to see the scatter. Like seeing bullet holes on a target without seeing the bull's-eye, a noise audit reveals judgment variability, giving us a practical path to improving judgment quality.
第5章
The Mathematics of Error: Why Noise Matters as Much as Bias
Bias and noise both contribute significantly to judgment error, and they play equal roles in calculating overall error. While bias involves consistent deviation in one direction, noise represents random variability in judgments. The impact of reducing noise by a certain amount equals the impact of reducing bias by the same amount.
The relationship between bias and noise can be expressed in a mathematical equation showing that overall error equals the sum of the squares of bias and noise. This Pythagorean-like relationship means bias and noise contribute equally and independently to overall error. Reducing either by the same amount produces identical improvements in accuracy, even though reducing noise while bias remains unchanged can create the counterintuitive appearance of making forecasts "more precisely wrong."
Despite this equal importance, noise reduction has traditionally received less attention than bias reduction. This neglect stems partly from our cognitive preferences-bias provides a satisfying causal explanation for errors, while noise requires statistical thinking that doesn't come naturally to us. Like a figure-ground demonstration in psychology, bias stands out as the compelling figure while noise recedes into the background.
When examining judgments across multiple cases, we can break down system noise into its components. Level noise occurs when some judges are consistently more severe than others across all cases-like "hanging judges" versus "bleeding-heart judges." In the sentencing study, the standard deviation of judges' average sentencing severity was 2.4 years. This variability correlates with judges' backgrounds, attitudes toward sentencing goals, geographic location, and political ideology-factors unrelated to justice itself.
Beyond level noise, judges differ in how they rank different cases-what the authors call "pattern noise." While level noise reflects consistent severity across all cases, pattern noise reveals judges' idiosyncratic reactions to specific case features. For example, one judge might be harsher than average generally but more lenient toward white-collar criminals. In the sentencing study, pattern noise contributed approximately equally to system noise as level noise did.
A third component, occasion noise, represents variability in a single person's judgments over time. Like a basketball player who never throws free throws exactly the same way, professionals don't produce identical judgments when faced with identical facts on different occasions. This creates a "second lottery" beyond the system noise lottery that determines which professional handles a case.
Evidence suggests that while occasion noise is troubling, it's generally smaller than the differences between individuals. You're less consistent over time than you think, but you remain more similar to yourself yesterday than to another person today.
第6章
When Algorithms Outperform Experts: The Noise-Free Alternative
Many predictive judgments can be evaluated for accuracy, offering valuable insights about noise and bias. When professionals make predictions-like rating job candidates on leadership potential-they typically use "clinical judgment," examining information holistically and intuitively. However, research consistently shows this approach performs poorly compared to simple mechanical rules or algorithms.
In one study, psychologists' predictions of job performance achieved only a 0.15 correlation with actual outcomes (55% concordance), barely better than chance. Statistical methods combining the same information achieved a 0.32 correlation (60% concordance). This gap between clinical and mechanical prediction reveals a fundamental limitation: humans consistently overestimate their predictive abilities, falling victim to the "illusion of validity."
Paul Meehl's groundbreaking research demonstrated that simple mechanical rules consistently outperform human judgment across diverse domains-a finding confirmed by later studies showing statistical models winning or tying in 128 of 136 comparisons with clinical judgment. Even more surprising, Goldberg's research showed that statistical models of individual judges consistently outperform the judges themselves. Though these models are crude approximations that eliminate both subtle rules and pattern noise from human judgment, they achieve better predictive accuracy.
This suggests that the gains from subtle rules in human judgment are generally insufficient to compensate for the detrimental effects of noise. As one study dramatically demonstrated, even randomly weighted linear models applied consistently outperformed human experts in predicting job performance.
The key advantage all mechanical approaches share-from simple formulas to complex machine learning-is that they're noise-free. Robyn Dawes discovered that "improper linear models" with equal weights perform nearly as well as complex regression models and far better than clinical judgments. This counterintuitive finding occurs because equal-weight models avoid overfitting to random fluctuations in data.
Even simpler "frugal models" produce surprisingly good predictions. A 2020 study applied this principle to bail decisions, creating a model using just two inputs: defendant's age and number of past missed court dates. This simple model, requiring no computer, matched more complex statistical models and outperformed human judges.
Despite overwhelming evidence of their superiority, algorithms remain underutilized in professional judgments. Medical diagnosis, hiring decisions, movie production, and sports management still rely heavily on intuition over formulas. Professionals resist algorithmic approaches due to overconfidence in their judgment, concerns about dehumanization, and fears of abdicating responsibility.
Even the most sophisticated algorithms face fundamental limits in prediction due to objective ignorance-what cannot possibly be known at the time of judgment. Both intractable uncertainty (what cannot be known) and imperfect information (what could be known but isn't) create objective ignorance that fundamentally limits prediction accuracy.
第7章
The Psychology of Noise: How Our Minds Create Variability
What mental mechanisms create the variability of our judgments? The psychology of noise reveals several key patterns that explain why different people-or even the same person at different times-reach different conclusions from identical information.
First, we often substitute difficult questions with easier ones-a core principle of the heuristics and biases approach. When judging a CEO candidate's chances of success, we typically neglect base rates (that roughly 72% of CEOs remain after two years) in favor of matching his description to our image of a successful CEO. The availability heuristic similarly substitutes ease of recall for actual frequency judgments, explaining why recent plane crashes temporarily inflate our perception of aviation risks.
Second, we form coherent impressions quickly and are slow to change them. When presented with information sequentially (like "Intelligent, Persistent" followed by "Cunning, Unprincipled"), our initial positive impression creates a halo effect that distorts how we interpret later information. This excessive coherence means we jump to conclusions then stick to them, interpreting new evidence to fit our initial judgment.
Third, we're remarkably prone to finding patterns in randomness. When evaluating job candidates, interviewers interpret the same behaviors differently based on their initial impressions-as when two interviewers heard a candidate left a previous job due to "strategic disagreement with the CEO" but one saw this as evidence of integrity while the other saw inflexibility.
These psychological mechanisms can produce both statistical bias and noise. Substitution creates noise when people replace the same question with different easier questions. Prejudgments produce noise when judges have different biases. Excessive coherence generates noise when information arrives in random order, causing initial impressions to vary arbitrarily between judges.
Mood significantly influences our judgments in surprising ways. People in good moods are more susceptible to stereotypes, more gullible toward meaningless but profound-sounding statements, and three times more likely to make utilitarian moral choices like sacrificing one person to save five. Other factors creating occasion noise include fatigue (doctors prescribe more opioids late in the day), weather (admissions officers favor academic attributes on cloudy days), and decision sequence (judges are 19% less likely to grant asylum after approving two previous cases).
Even in tightly controlled environments, occasion noise persists mysteriously. A study of memory performance found that external factors like sleep, time of day, and practice explained only 11% of performance variation. The strongest predictor was how well subjects performed on the immediately preceding task, suggesting that performance ebbs and flows due to intrinsic neural variability.
Group decision-making adds another layer to the noise problem. Who speaks first, who appears confident, even who smiles at the right moment-all these irrelevant factors can send similar groups in dramatically different directions. When people announce judgments sequentially, early opinions disproportionately influence later ones, creating informational cascades that can lead entire groups in potentially misguided directions based on initial speakers' views.
第8章
Decision Hygiene: Practical Strategies for Reducing Noise
How can organizations improve professional judgments and reduce noise? The authors propose a comprehensive approach called "decision hygiene"-preventive measures that reduce error before it occurs. Like handwashing in medicine, decision hygiene prevents a range of problems even though you'll never know which specific errors you prevented.
The first strategy involves selecting better judges. Some people consistently perform better than others at judgment tasks, producing results that are both less noisy and less biased. Three factors determine judgment quality: what you know (training and experience), how well you think (intelligence), and how you think (cognitive style). The most predictive measure for forecasting performance is "actively open-minded thinking"-the willingness to search for information contradicting one's preexisting hypotheses.
The second strategy involves properly sequencing information to limit premature intuitions. In forensic science, cognitive neuroscientist Itiel Dror found that when fingerprint examiners were exposed to biasing contextual information (like "the suspect confessed"), their judgments often changed. His "linear sequential unmasking" approach requires examiners to document their judgments at each step before seeing potentially biasing information. Experts should analyze latent prints before seeing exemplars, record judgments before accessing contextual information, and document any subsequent changes.
The third strategy combines selection and aggregation in forecasting. The Good Judgment Project, founded by Philip Tetlock, Barbara Mellers, and Don Moore, found that averaging multiple forecasts mathematically reduces noise by dividing it by the square root of the number of judgments averaged-100 judgments reduces noise by 90%. Their research identified "superforecasters" who excel at breaking down complex problems into components, asking "What would it take for the answer to be yes or no?" rather than relying on gut feelings. These individuals prioritize base rates and embody "perpetual beta"-continuous self-improvement through research, self-criticism, and synthesizing diverse perspectives.
The fourth strategy uses guidelines to decompose complex decisions into simpler judgments on predefined dimensions. The Apgar score exemplifies this approach by evaluating newborns on five key measures, each scored 0-2. Guidelines work by focusing clinicians on empirically important predictors, simplifying individual judgments, and specifying how to weight components. Similar successful approaches include the Centor score for strep throat diagnosis and the BI-RADS system for mammogram interpretation.
The fifth strategy involves structured interviews in hiring. Google discovered their recruiting interviews had "zero relationship" with performance and implemented evidence-based improvements. They adopted structured judgment with three key principles: decomposition (breaking evaluation into specific components), independence (collecting information separately for each component through structured behavioral interviews), and delayed holistic judgment (making final decisions only after systematically gathering all evidence).
Finally, the mediating assessments protocol applies these principles broadly to organizational decision-making. The approach treats options like job candidates that require structured evaluation across multiple dimensions. It implements several decision hygiene techniques: structuring decisions into independent assessments, using outside-view reference points, sequencing information properly, and aggregating independent judgments.
第9章
Finding the Right Balance: When to Embrace Noise
Despite the compelling case for noise reduction, many people resist such efforts, as evidenced by the negative judicial reaction to sentencing guidelines. Some view rules as rigid, dehumanizing, and unfair in their own way. Critics argue that mechanical solutions cannot satisfy "the demands of justice" and that focusing on rules reflects "a fear of judging."
Noise reduction strategies often face objections that they're too expensive or impractical. A high school teacher grading essays might find that adding a second reader or using structured assessment tools improves accuracy but requires unaffordable time investments. Similarly, hospitals might identify that diagnostic variability could be reduced through additional testing, but the tests themselves might be invasive, dangerous, and costly.
When people are denied opportunities through rigid rules or algorithms rather than individualized human judgment, they often object on grounds of dignity. Many insist on face-to-face interaction where a human exercises discretion and considers their unique circumstances. This preference for case-by-case judgment has deep moral foundations across cultures, politics, law, theology, and literature.
Clear rules that eliminate noise may create opportunities for gaming the system. The tax code illustrates this dilemma-while identical taxpayers shouldn't be treated differently, eliminating all noise would enable clever taxpayers to find loopholes. Similarly, organizations that specifically list prohibited behaviors might inadvertently permit harmful conduct not explicitly covered.
Noise reduction efforts might also squelch motivation, creativity, and engagement. People in positions of authority resist having their discretion removed, feeling diminished and constrained. When employees can respond to situations in their own way, they enjoy their jobs more and may develop fresh ideas.
Organizations must choose between rules (which eliminate discretion) and standards (which grant it). Rules reduce noise by answering factual questions, while standards require judges to interpret open-ended terms, inevitably producing noise. When organizations are sharply divided or lack sufficient information, standards may be easier to implement than rules. Leaders might agree on broad principles without agreeing on specifics.
The choice between rules and standards should depend on two key factors: decision costs and error costs. Standards impose higher decision costs on judges who must spend time giving them content, while rules allow for faster, more straightforward decisions. However, creating good rules initially requires significant effort. Error costs depend on the number and magnitude of mistakes. When agents are knowledgeable, reliable, and practice decision hygiene, standards may work well with minimal noise. Rules become necessary when agents can't be fully trusted.
第10章
Toward a Less Noisy World
Imagine organizations redesigned to minimize noise: hospitals, hiring committees, forecasters, government agencies, insurance companies, and justice systems would routinely conduct noise audits. Leaders would deploy algorithms to replace or supplement human judgment. Complex decisions would be broken into simpler mediating assessments. Decision hygiene would be standard practice.
Independent judgments would be elicited before discussion and then aggregated. Meetings would become more structured, with outside views systematically integrated and disagreements more constructively resolved. This less noisy world would save money, improve public safety and health, increase fairness, and prevent countless avoidable errors.
While noise can never be completely eliminated-and shouldn't be in some contexts-the evidence is clear that most organizations have far more noise than they realize or would consider acceptable. The first step toward improvement is recognition-understanding that noise exists and measuring its magnitude through noise audits.
The fundamental problem remains that without noise audits, organizations remain unaware of how much noise exists in their judgments, making cost-benefit calculations impossible. Noise is a hidden epidemic affecting virtually every domain where human judgment matters. By taking noise seriously and implementing decision hygiene practices, we can create fairer, more accurate, and more consistent judgments-an opportunity worth seizing.