Chapter 4
Your Mind Is a Measuring Instrument
When making judgments, we're essentially using our minds as measuring instruments. Just as a thermometer assigns temperature values to objects, our judgments assign values on various scales-whether sentencing criminals, pricing insurance, or diagnosing illnesses. Like all measurements, judgments contain error-some bias, some noise-and perfect accuracy remains unattainable.
Professional judgments occupy a middle ground between factual questions (where reasonable people should agree perfectly) and matters of taste (where disagreement is acceptable). They're characterized by an "expectation of bounded disagreement"-reasonable people might disagree, but not wildly so. The acceptable range depends on the problem's difficulty, with agreement typically easier on absurd judgments than on reasonable ones.
Making a judgment involves several cognitive processes that introduce variability. Consider evaluating a CEO candidate: we selectively attend to certain details while ignoring others-some focus on his Harvard background, others on his abrasive style. We then informally integrate these cues into an overall impression without following a structured process. Finally, we convert this impression into a specific number on a scale, often through an intuitive process where numbers "feel right" rather than through explicit calculation. Each step introduces noise into our judgments, explaining why MBA students' estimates of the same candidate's success probability ranged from 10% to 95%.
When making judgments, we're not actually trying to match an external outcome (which often doesn't exist yet), but rather seeking an internal signal of coherence-that feeling when our answer feels right because it comfortably fits with the evidence. This "internal signal of judgment completion" occurs whether the judgment is verifiable or not.
Noise in judgments indicates problems in different ways. In predictive judgments (like medical diagnoses), disagreement means at least one person must be wrong, potentially causing serious consequences. In evaluative judgments (like sentencing), noise violates expectations of fairness and consistency-turning justice into a lottery. Even when unfairness isn't the primary concern, system noise represents inconsistency that damages credibility, as when similar customer complaints receive dramatically different responses.
The good news is that we can measure noise without knowing the "true" value of a judgment-we simply need multiple judgments of the same problem to see the scatter. Like seeing bullet holes on a target without seeing the bull's-eye, a noise audit reveals judgment variability, giving us a practical path to improving judgment quality.
Chapter 5
The Mathematics of Error: Why Noise Matters as Much as Bias
Bias and noise both contribute significantly to judgment error, and they play equal roles in calculating overall error. While bias involves consistent deviation in one direction, noise represents random variability in judgments. The impact of reducing noise by a certain amount equals the impact of reducing bias by the same amount.
The relationship between bias and noise can be expressed in a mathematical equation showing that overall error equals the sum of the squares of bias and noise. This Pythagorean-like relationship means bias and noise contribute equally and independently to overall error. Reducing either by the same amount produces identical improvements in accuracy, even though reducing noise while bias remains unchanged can create the counterintuitive appearance of making forecasts "more precisely wrong."
Despite this equal importance, noise reduction has traditionally received less attention than bias reduction. This neglect stems partly from our cognitive preferences-bias provides a satisfying causal explanation for errors, while noise requires statistical thinking that doesn't come naturally to us. Like a figure-ground demonstration in psychology, bias stands out as the compelling figure while noise recedes into the background.
When examining judgments across multiple cases, we can break down system noise into its components. Level noise occurs when some judges are consistently more severe than others across all cases-like "hanging judges" versus "bleeding-heart judges." In the sentencing study, the standard deviation of judges' average sentencing severity was 2.4 years. This variability correlates with judges' backgrounds, attitudes toward sentencing goals, geographic location, and political ideology-factors unrelated to justice itself.
Beyond level noise, judges differ in how they rank different cases-what the authors call "pattern noise." While level noise reflects consistent severity across all cases, pattern noise reveals judges' idiosyncratic reactions to specific case features. For example, one judge might be harsher than average generally but more lenient toward white-collar criminals. In the sentencing study, pattern noise contributed approximately equally to system noise as level noise did.
A third component, occasion noise, represents variability in a single person's judgments over time. Like a basketball player who never throws free throws exactly the same way, professionals don't produce identical judgments when faced with identical facts on different occasions. This creates a "second lottery" beyond the system noise lottery that determines which professional handles a case.
Evidence suggests that while occasion noise is troubling, it's generally smaller than the differences between individuals. You're less consistent over time than you think, but you remain more similar to yourself yesterday than to another person today.
Chapter 6
When Algorithms Outperform Experts: The Noise-Free Alternative
Many predictive judgments can be evaluated for accuracy, offering valuable insights about noise and bias. When professionals make predictions-like rating job candidates on leadership potential-they typically use "clinical judgment," examining information holistically and intuitively. However, research consistently shows this approach performs poorly compared to simple mechanical rules or algorithms.
In one study, psychologists' predictions of job performance achieved only a 0.15 correlation with actual outcomes (55% concordance), barely better than chance. Statistical methods combining the same information achieved a 0.32 correlation (60% concordance). This gap between clinical and mechanical prediction reveals a fundamental limitation: humans consistently overestimate their predictive abilities, falling victim to the "illusion of validity."
Paul Meehl's groundbreaking research demonstrated that simple mechanical rules consistently outperform human judgment across diverse domains-a finding confirmed by later studies showing statistical models winning or tying in 128 of 136 comparisons with clinical judgment. Even more surprising, Goldberg's research showed that statistical models of individual judges consistently outperform the judges themselves. Though these models are crude approximations that eliminate both subtle rules and pattern noise from human judgment, they achieve better predictive accuracy.
This suggests that the gains from subtle rules in human judgment are generally insufficient to compensate for the detrimental effects of noise. As one study dramatically demonstrated, even randomly weighted linear models applied consistently outperformed human experts in predicting job performance.
The key advantage all mechanical approaches share-from simple formulas to complex machine learning-is that they're noise-free. Robyn Dawes discovered that "improper linear models" with equal weights perform nearly as well as complex regression models and far better than clinical judgments. This counterintuitive finding occurs because equal-weight models avoid overfitting to random fluctuations in data.
Even simpler "frugal models" produce surprisingly good predictions. A 2020 study applied this principle to bail decisions, creating a model using just two inputs: defendant's age and number of past missed court dates. This simple model, requiring no computer, matched more complex statistical models and outperformed human judges.
Despite overwhelming evidence of their superiority, algorithms remain underutilized in professional judgments. Medical diagnosis, hiring decisions, movie production, and sports management still rely heavily on intuition over formulas. Professionals resist algorithmic approaches due to overconfidence in their judgment, concerns about dehumanization, and fears of abdicating responsibility.
Even the most sophisticated algorithms face fundamental limits in prediction due to objective ignorance-what cannot possibly be known at the time of judgment. Both intractable uncertainty (what cannot be known) and imperfect information (what could be known but isn't) create objective ignorance that fundamentally limits prediction accuracy.
Chapter 7
The Psychology of Noise: How Our Minds Create Variability
What mental mechanisms create the variability of our judgments? The psychology of noise reveals several key patterns that explain why different people-or even the same person at different times-reach different conclusions from identical information.
First, we often substitute difficult questions with easier ones-a core principle of the heuristics and biases approach. When judging a CEO candidate's chances of success, we typically neglect base rates (that roughly 72% of CEOs remain after two years) in favor of matching his description to our image of a successful CEO. The availability heuristic similarly substitutes ease of recall for actual frequency judgments, explaining why recent plane crashes temporarily inflate our perception of aviation risks.
Second, we form coherent impressions quickly and are slow to change them. When presented with information sequentially (like "Intelligent, Persistent" followed by "Cunning, Unprincipled"), our initial positive impression creates a halo effect that distorts how we interpret later information. This excessive coherence means we jump to conclusions then stick to them, interpreting new evidence to fit our initial judgment.
Third, we're remarkably prone to finding patterns in randomness. When evaluating job candidates, interviewers interpret the same behaviors differently based on their initial impressions-as when two interviewers heard a candidate left a previous job due to "strategic disagreement with the CEO" but one saw this as evidence of integrity while the other saw inflexibility.
These psychological mechanisms can produce both statistical bias and noise. Substitution creates noise when people replace the same question with different easier questions. Prejudgments produce noise when judges have different biases. Excessive coherence generates noise when information arrives in random order, causing initial impressions to vary arbitrarily between judges.
Mood significantly influences our judgments in surprising ways. People in good moods are more susceptible to stereotypes, more gullible toward meaningless but profound-sounding statements, and three times more likely to make utilitarian moral choices like sacrificing one person to save five. Other factors creating occasion noise include fatigue (doctors prescribe more opioids late in the day), weather (admissions officers favor academic attributes on cloudy days), and decision sequence (judges are 19% less likely to grant asylum after approving two previous cases).
Even in tightly controlled environments, occasion noise persists mysteriously. A study of memory performance found that external factors like sleep, time of day, and practice explained only 11% of performance variation. The strongest predictor was how well subjects performed on the immediately preceding task, suggesting that performance ebbs and flows due to intrinsic neural variability.
Group decision-making adds another layer to the noise problem. Who speaks first, who appears confident, even who smiles at the right moment-all these irrelevant factors can send similar groups in dramatically different directions. When people announce judgments sequentially, early opinions disproportionately influence later ones, creating informational cascades that can lead entire groups in potentially misguided directions based on initial speakers' views.
Chapter 8
Decision Hygiene: Practical Strategies for Reducing Noise
How can organizations improve professional judgments and reduce noise? The authors propose a comprehensive approach called "decision hygiene"-preventive measures that reduce error before it occurs. Like handwashing in medicine, decision hygiene prevents a range of problems even though you'll never know which specific errors you prevented.
The first strategy involves selecting better judges. Some people consistently perform better than others at judgment tasks, producing results that are both less noisy and less biased. Three factors determine judgment quality: what you know (training and experience), how well you think (intelligence), and how you think (cognitive style). The most predictive measure for forecasting performance is "actively open-minded thinking"-the willingness to search for information contradicting one's preexisting hypotheses.
The second strategy involves properly sequencing information to limit premature intuitions. In forensic science, cognitive neuroscientist Itiel Dror found that when fingerprint examiners were exposed to biasing contextual information (like "the suspect confessed"), their judgments often changed. His "linear sequential unmasking" approach requires examiners to document their judgments at each step before seeing potentially biasing information. Experts should analyze latent prints before seeing exemplars, record judgments before accessing contextual information, and document any subsequent changes.
The third strategy combines selection and aggregation in forecasting. The Good Judgment Project, founded by Philip Tetlock, Barbara Mellers, and Don Moore, found that averaging multiple forecasts mathematically reduces noise by dividing it by the square root of the number of judgments averaged-100 judgments reduces noise by 90%. Their research identified "superforecasters" who excel at breaking down complex problems into components, asking "What would it take for the answer to be yes or no?" rather than relying on gut feelings. These individuals prioritize base rates and embody "perpetual beta"-continuous self-improvement through research, self-criticism, and synthesizing diverse perspectives.
The fourth strategy uses guidelines to decompose complex decisions into simpler judgments on predefined dimensions. The Apgar score exemplifies this approach by evaluating newborns on five key measures, each scored 0-2. Guidelines work by focusing clinicians on empirically important predictors, simplifying individual judgments, and specifying how to weight components. Similar successful approaches include the Centor score for strep throat diagnosis and the BI-RADS system for mammogram interpretation.
The fifth strategy involves structured interviews in hiring. Google discovered their recruiting interviews had "zero relationship" with performance and implemented evidence-based improvements. They adopted structured judgment with three key principles: decomposition (breaking evaluation into specific components), independence (collecting information separately for each component through structured behavioral interviews), and delayed holistic judgment (making final decisions only after systematically gathering all evidence).
Finally, the mediating assessments protocol applies these principles broadly to organizational decision-making. The approach treats options like job candidates that require structured evaluation across multiple dimensions. It implements several decision hygiene techniques: structuring decisions into independent assessments, using outside-view reference points, sequencing information properly, and aggregating independent judgments.
Chapter 9
Finding the Right Balance: When to Embrace Noise
Despite the compelling case for noise reduction, many people resist such efforts, as evidenced by the negative judicial reaction to sentencing guidelines. Some view rules as rigid, dehumanizing, and unfair in their own way. Critics argue that mechanical solutions cannot satisfy "the demands of justice" and that focusing on rules reflects "a fear of judging."
Noise reduction strategies often face objections that they're too expensive or impractical. A high school teacher grading essays might find that adding a second reader or using structured assessment tools improves accuracy but requires unaffordable time investments. Similarly, hospitals might identify that diagnostic variability could be reduced through additional testing, but the tests themselves might be invasive, dangerous, and costly.
When people are denied opportunities through rigid rules or algorithms rather than individualized human judgment, they often object on grounds of dignity. Many insist on face-to-face interaction where a human exercises discretion and considers their unique circumstances. This preference for case-by-case judgment has deep moral foundations across cultures, politics, law, theology, and literature.
Clear rules that eliminate noise may create opportunities for gaming the system. The tax code illustrates this dilemma-while identical taxpayers shouldn't be treated differently, eliminating all noise would enable clever taxpayers to find loopholes. Similarly, organizations that specifically list prohibited behaviors might inadvertently permit harmful conduct not explicitly covered.
Noise reduction efforts might also squelch motivation, creativity, and engagement. People in positions of authority resist having their discretion removed, feeling diminished and constrained. When employees can respond to situations in their own way, they enjoy their jobs more and may develop fresh ideas.
Organizations must choose between rules (which eliminate discretion) and standards (which grant it). Rules reduce noise by answering factual questions, while standards require judges to interpret open-ended terms, inevitably producing noise. When organizations are sharply divided or lack sufficient information, standards may be easier to implement than rules. Leaders might agree on broad principles without agreeing on specifics.
The choice between rules and standards should depend on two key factors: decision costs and error costs. Standards impose higher decision costs on judges who must spend time giving them content, while rules allow for faster, more straightforward decisions. However, creating good rules initially requires significant effort. Error costs depend on the number and magnitude of mistakes. When agents are knowledgeable, reliable, and practice decision hygiene, standards may work well with minimal noise. Rules become necessary when agents can't be fully trusted.
Chapter 10
Toward a Less Noisy World
Imagine organizations redesigned to minimize noise: hospitals, hiring committees, forecasters, government agencies, insurance companies, and justice systems would routinely conduct noise audits. Leaders would deploy algorithms to replace or supplement human judgment. Complex decisions would be broken into simpler mediating assessments. Decision hygiene would be standard practice.
Independent judgments would be elicited before discussion and then aggregated. Meetings would become more structured, with outside views systematically integrated and disagreements more constructively resolved. This less noisy world would save money, improve public safety and health, increase fairness, and prevent countless avoidable errors.
While noise can never be completely eliminated-and shouldn't be in some contexts-the evidence is clear that most organizations have far more noise than they realize or would consider acceptable. The first step toward improvement is recognition-understanding that noise exists and measuring its magnitude through noise audits.
The fundamental problem remains that without noise audits, organizations remain unaware of how much noise exists in their judgments, making cost-benefit calculations impossible. Noise is a hidden epidemic affecting virtually every domain where human judgment matters. By taking noise seriously and implementing decision hygiene practices, we can create fairer, more accurate, and more consistent judgments-an opportunity worth seizing.