Capítulo 1
The Crystal Ball Revolution: How Ordinary People Outpredict Experts
In 2011, a retired USDA employee from Nebraska named Bill Flack began answering difficult geopolitical forecasting questions like "Will Russia annex additional Ukrainian territory?" Despite lacking the celebrity status of pundits like Tom Friedman, Bill consistently makes remarkably accurate predictions. He's part of a group called "superforecasters" - ordinary people who consistently outperform experts, intelligence analysts, and prediction markets when forecasting global events.
This phenomenon emerged from the Good Judgment Project, a forecasting tournament sponsored by the Intelligence Advanced Research Projects Activity (IARPA) where over 20,000 volunteers predicted global events ranging from Russian protests to gold prices. The project revealed two critical findings: first, foresight is real and measurable; second, forecasting excellence comes not from innate talent but from specific thinking habits that anyone can learn.
My research comparing expert forecasters to "dart-throwing chimpanzees" has been widely misinterpreted. What I actually found was that average experts performed only slightly better than random guessing on political and economic questions. This doesn't mean all expertise is worthless - some experts significantly outperform others. I position myself as an "optimistic skeptic" - recognizing forecasting's limits while believing intelligent, open-minded people can develop meaningful predictive skills.
While acknowledging prediction's limits in complex systems (the famous "butterfly effect"), dismissing all forecasting as futile would be a mistake. Our daily lives are filled with reliable predictions - from rush hour traffic patterns to sunrise times. Reality contains both clocklike predictability and cloudlike uncertainty, with the balance depending on what we're predicting, timeframe, and circumstances.
Capítulo 2
The Deadly Certainty Trap
Our minds rush to judgment and resist changing course, even when wrong. This cognitive trap is illustrated by Archie Cochrane's story from 1956, when a renowned specialist confidently diagnosed him with terminal cancer. Cochrane accepted this death sentence without question, methodically planning his remaining time - only to discover the specialist was completely wrong. Neither man had questioned the diagnosis or waited for the pathologist's report.
Medical history reveals a disturbing pattern: most treatments throughout history were useless or harmful. From ancient Egyptian ostrich egg poultices to George Washington's physicians who bled, purged, and blistered him to death in 1799, physicians rarely questioned their ineffective methods. As surgeon-historian Ira Rutkow observed, these physicians were "like blind men arguing over the colors of the rainbow."
The cure for medicine's plague of certainty was painfully slow to emerge. Though James Lind conducted what resembled a modern experiment in 1747 by testing different scurvy treatments on sailors, not until the twentieth century did randomized trials and statistical methods take hold. Archie Cochrane, who championed scientific validation in the 1950s and 1960s, faced fierce opposition from physicians with what he called "the God complex" - the belief that their judgment alone revealed truth.
Our thinking operates in two distinct systems. System 1 is fast, automatic, and constantly running in the background - delivering instant judgments without conscious effort. System 2 is our familiar conscious thought, which requires focus and effort. Most of us rely heavily on System 1's snap judgments, which follow primitive psycho-logic: if it feels true, it is.
We often unconsciously substitute hard questions with easier ones - a mental sleight-of-hand called attribute substitution or "bait and switch." When Cochrane was told he had terminal cancer, he unconsciously substituted "Is this the sort of person who should know if I have cancer?" for "Do I have cancer?" Our minds crave coherent explanations, leading us to confirmation bias - eagerly gathering supportive evidence while dismissing contradictory information.
Capítulo 3
The Science of Prediction
When physicians learned to doubt themselves, they turned to randomized controlled trials for scientific testing. While measuring forecasting accuracy might seem straightforward, it's surprisingly complex. Consider Steve Ballmer's infamous 2007 prediction that "There's no chance that the iPhone is going to get any significant market share." Though widely mocked as spectacularly wrong, his full statement was more nuanced, referring to the global mobile phone market. By 2013, iPhones had about 6% of that market - higher than his predicted "2% or 3%" but not laugh-out-loud wrong.
Forecasts must be testable to be valuable, but many obstacles stand in the way. Forecasts often lack clear timelines or precise definitions of key terms, making them impossible to verify. Probability presents an even bigger challenge - we can't rerun history to determine if a 70% chance prediction was accurate. Sherman Kent, a CIA legend, tried to solve this by creating a numerical probability scale to replace vague language, but faced resistance from those who preferred ambiguous terms that couldn't be proven wrong.
To properly evaluate forecasts, we need benchmarks and comparability. Comparing forecasters requires a level playing field - predicting weather in highly variable Springfield is harder than in stable Phoenix. This means carefully crafted experiments are essential: assembling forecasters, asking precise questions with clear timeframes, requiring numerical probabilities, and waiting for results.
In the mid-1980s, I began a forecasting experiment that would span decades. I recruited 284 serious professionals who analyzed political and economic trends - academics, government officials, and others. They made roughly 28,000 forecasts on diverse topics across one to ten-year timeframes. This allowed comparison between true subject-matter experts and well-informed generalists.
The average expert performed no better than random guessing - like a dart-throwing chimpanzee. However, two distinct groups emerged: one that failed to beat random chance, and another that modestly outperformed it. The difference wasn't in credentials or information access, but thinking style. The underperformers, whom I dubbed "hedgehogs," organized thinking around Big Ideas and ideologies, expressed high confidence, and resisted changing their minds when wrong. The better performers - "foxes" - used diverse analytical tools, acknowledged uncertainty, spoke of probabilities rather than certainties, and more readily admitted errors.
Capítulo 4
The Birth of Superforecasting
The 2002 National Intelligence Estimate that concluded Iraq possessed weapons of mass destruction represents one of history's worst intelligence failures. The critical error wasn't being wrong but expressing absolute certainty where uncertainty existed. The intelligence community never seriously explored alternative possibilities or conducted "red team" analyses challenging prevailing views.
Despite spending $50 billion annually and employing 100,000 people (including 20,000 analysts), the intelligence community has never systematically tracked forecast accuracy. Officials may claim 80-90% accuracy rates, but these are mere guesses. Instead of measuring results, the IC holds analysts accountable only for process - checking whether they followed proper procedures rather than whether their judgments proved correct.
After presenting superforecasters' success, I urge skepticism. Consider my coin-toss thought experiment: if 2,800 volunteers predicted 104 coin tosses, pure chance would create a bell curve with some appearing extraordinarily skilled through random luck. To determine if superforecasters possess genuine skill or just luck, we must examine regression to the mean - if purely lucky, their performance should regress completely to average in subsequent years.
Remarkably, superforecasters didn't regress; they actually increased their lead in years 2 and 3. While approximately 30% fall from the top ranks each year, 70% consistently remain superforecasters - odds of 1 in 3 given the 0.65 year-to-year correlation, versus 1 in 100,000,000 if purely random. This confirms superforecasters possess genuine skill, though luck still plays a role. The big question remains: why are they so good?
Capítulo 5
The Superforecaster Mindset
In 2008, Sanford "Sandy" Sillman was diagnosed with multiple sclerosis. Though not life-threatening, the disease weakened him, caused pain, and made walking and typing difficult. By 2011, at fifty-seven, he could see he would soon need to leave his job as an atmospheric scientist. Anticipating the void this would create, Sandy joined the Good Judgment Project as a forecaster, seeking a meaningful transition activity that would "keep his mind alive" without the pressure of his former work.
The ability to break down complex questions into manageable parts is crucial for superforecasters. Enrico Fermi famously challenged students to estimate seemingly impossible questions like "How many piano tuners are in Chicago?" The key is decomposing the question into smaller, more answerable parts: how many pianos exist in Chicago, how often they're tuned, how long tuning takes, and how many hours tuners work annually. Even with imperfect information, this structured approach produces surprisingly accurate estimates.
When IARPA asked forecasters if Swiss or French inquiries would find elevated polonium levels in Yasser Arafat's exhumed body, the question demanded careful analysis beyond gut reactions. Most people would instinctively answer a different question-"Did Israel poison Arafat?"-rather than the actual question posed. This is System 1's classic bait-and-switch.
Superforecaster Bill Flack avoided this trap by breaking down the question. Though lacking Middle East expertise, he first examined whether polonium could still be detected years after death. He then identified multiple pathways to contamination and noted that only one of two European teams needed a positive result for a "yes" answer. This methodical breakdown created a roadmap for analysis that avoided the intuitive error most would make.
Before diving into case specifics, superforecasters establish a "base rate" or "outside view"-how common something is within a broader class. While the inside view (specific details) is concrete and engaging, the outside view provides a statistical foundation. For example, when predicting pet ownership for a family, starting with the fact that 62% of American households own pets creates a better starting point than analyzing the family's specific characteristics.
After establishing the outside view, superforecasters explore the inside view through targeted investigation. Using a framework, Bill Flack examined specific hypotheses about Arafat's possible polonium poisoning: Israel could have poisoned him; Palestinian enemies could have done it; or someone could have contaminated his remains posthumously. For each hypothesis, he identified necessary conditions, methodically analyzing each one.
The final step is merging outside and inside views. Superforecaster David Rogg demonstrated this when addressing whether Islamist militants would attack certain European countries between January and March 2015. First, he established the outside view by counting previous attacks in those countries (six in five years, or 1.2 per year). Then he adjusted for inside-view factors, calculated the proportion of the year in question, and multiplied to reach a 34% probability.
Capítulo 6
The Numbers Game
We live in the Big Data era where data scientists use powerful computers and complex math to extract meaning from vast information networks. Superforecasters are comfortable with numbers - most aced basic numeracy tests and many have backgrounds in math, science or computer programming. While some occasionally deploy mathematical models (like Bill Flack using Monte Carlo simulations for currency forecasts), most superforecasters don't rely on complex mathematical techniques. Their numeracy helps in more subtle ways than performing statistical wizardry.
In early 2011, intelligence analysts had to judge whether Osama bin Laden was hiding in a peculiar Abbottabad compound. The movie Zero Dark Thirty dramatizes this scenario, showing CIA Director Leon Panetta (played by James Gandolfini) asking analysts for their confidence levels. The fictional Panetta grows frustrated when analysts offer varying probabilities, calling it a "clusterfuck." But this diversity of independent judgments actually represents valuable "wisdom of the crowd." The real Leon Panetta welcomed the diversity of opinions ranging from 30-40% to over 90%, explaining, "I encourage the people around me not to tell me what they thought I wanted to hear but what they believed."
Mark Bowden's book The Finish describes a similar scene with President Obama receiving varied CIA probability estimates about bin Laden's presence - from 95% to as low as 30-40%. Obama responded by declaring it "fifty-fifty" and "a flip of the coin." As Amos Tversky once joked, most people only have three probability settings: "gonna happen," "not gonna happen," and "maybe" - a simplification that captures a fundamental truth about human judgment.
Humans coped with uncertainty long before formal probability theory emerged. Our ancestors relied on intuitive judgments with essentially three settings: yes, no, or maybe. This simplified approach made evolutionary sense - when spotting movement in tall grass, you needed quick decisions about lions, not fine-grained probability estimates. Research shows we overvalue certainty, willing to pay significantly more to reduce risk from 5% to 0% than from 10% to 5%. We also equate confidence with competence, distrusting forecasters who express middling probabilities.
Scientists approach probability radically differently, embracing uncertainty as an ineradicable element of reality. Scientific "facts" once considered solid can be overturned by new evidence - all knowledge remains tentative. This means the intuitive two- and three-setting mental dials (yes/no/maybe) are fundamentally flawed. Instead, probabilistic thinking requires subdividing "maybe" into precise numerical gradations.
Superforecasters, like scientists and mathematicians, are natural probabilistic thinkers who grasp the distinction between "epistemic" uncertainty (theoretically knowable) and "aleatory" uncertainty (fundamentally unknowable). When facing questions loaded with irreducible uncertainty, they remain cautious, keeping initial estimates between 35% and 65%. Unlike average forecasters who frequently use "fifty-fifty" as a stand-in for "maybe," superforecasters are remarkably precise, using granular estimates like 37% rather than rounded numbers like 30% or 40%.
Capítulo 7
The Continuous Learning Cycle
Superforecasting follows a methodical process: unpacking questions into components, scrutinizing assumptions, adopting both outside and inside views, exploring different perspectives, synthesizing information like a dragonfly's compound vision, and expressing judgments on fine-grained probability scales. But this initial forecast is just the beginning. Unlike lottery tickets, forecasts must be continuously updated as new information emerges.
Updating forecasts requires distinguishing between obvious developments that everyone sees and subtler information that provides competitive advantage. Bill Flack demonstrated this skill when the Swiss team investigating Yasser Arafat's remains announced a delay for additional testing. Understanding polonium's properties, Bill inferred this likely meant they had detected polonium and were testing to rule out natural lead decay as its source. He cautiously raised his forecast to 65% yes before others recognized the significance - a decision validated when polonium was indeed found.
Underreaction to new information can destroy forecasts. Sometimes it's simply due to practical constraints - Joshua Frankel failed to update his Syria forecast when swamped with work. More insidious is when psychological biases interfere. The most tenacious form of underreaction comes from belief perseverance - our reluctance to change established beliefs. The more central a belief is to our identity, the harder it is to dislodge, like removing a critical block from a Jenga tower.
Overreaction to new information can be just as damaging. Psychology experiments show we're swayed by completely irrelevant information - a phenomenon called the dilution effect. This bias explains much of the excessive trading in financial markets. As Burton Malkiel observed, many investors hop between stocks "as if selecting and discarding cards in gin rummy," with studies showing frequent traders earned 11.4% annual returns during a period when the market averaged 17.9%.
Tim Minto, a Vancouver software engineer who won the third season of the IARPA tournament with an impressive 0.15 Brier score, exemplifies masterful belief updating. Unlike other forecasters, Tim's approach involves constant but tiny adjustments - averaging just 3.5% changes per update. On a question about Syrian refugees, he updated 34 times over three months, gradually steering toward the correct answer without dramatic swings. This methodical approach helps navigate between the twin dangers of forecasting: underreaction and overreaction.
Capítulo 8
The Growth Mindset
Failure drove Mary Simpson, a PhD economist, to become a superforecaster. Despite her expertise, she completely missed the 2007 financial crisis until it was too late, watching her retirement savings crater. This frustrating experience motivated her to improve her forecasting abilities, believing she could and should do better. Psychologist Carol Dweck would identify this as a "growth mindset" - the belief that abilities can be developed through effort and learning, rather than being fixed traits.
John Maynard Keynes exemplified this growth mindset. Beyond his macroeconomic theories, Keynes was a remarkably successful investor who managed money between the World Wars. Despite devastating setbacks - nearly wiped out by wrong currency forecasts in 1920 and blindsided by the 1929 crash - Keynes rebounded stronger each time by embracing his mistakes as learning opportunities. He prided himself on his willingness to change his mind, famously quipping "There is no harm in being sometimes wrong, especially if one is promptly found out."
Learning to forecast requires actually forecasting, just as learning to ride a bicycle requires getting on one. Philosopher Michael Polanyi demonstrated this by writing a technically perfect explanation of bicycle physics that would leave anyone who read it still unable to ride. We need "tacit knowledge" that can only come from direct experience.
Effective practice requires clear, timely feedback - something most forecasters lack. Without it, forecasters suffer from hindsight bias, the tendency to misremember predictions once outcomes are known. In Fischhoff's experiments, people consistently recalled their estimates as closer to actual outcomes than they really were. When experts were asked in 1988 about the likelihood of the Communist Party losing power in the Soviet Union, then asked in 1992-93 to recall their estimates, they remembered probabilities 31 percentage points higher than their actual forecasts - a classic "I knew it all along" effect.
Superforecasters conduct thorough postmortems of their predictions, eager to learn from both successes and failures. After a question closes, they analyze what went wrong and how to improve, sharing lengthy discussions with teammates. Unlike most experts who dismiss "I was almost wrong" scenarios while embracing "I was almost right" ones, superforecasters readily acknowledge when luck played a role in their successful predictions.
Forecasting requires what Angela Duckworth calls "grit" - passionate perseverance toward long-term goals despite frustration and failure. Elizabeth Sloane exemplifies this, volunteering for the Good Judgment Project while battling brain cancer to "re-grow her synapses." Anne Kilkenny, a housewife from Alaska with no geopolitical background, demonstrated remarkable tenacity by researching UN refugee data in the Central African Republic, even emailing the agency directly and overcoming language barriers to get better information.
Capítulo 9
The Power of Collaboration
The Kennedy administration's contrasting experiences with the Bay of Pigs disaster and the Cuban missile crisis illustrate how the same team can produce both catastrophic failures and brilliant successes. After the Bay of Pigs fiasco, Kennedy transformed his decision-making culture by instituting skepticism, designating "intellectual watchdogs" to challenge assumptions, setting aside protocol and hierarchy, bringing in fresh perspectives, and sometimes absenting himself to encourage freer discussion.
The first year's results of our forecasting tournament were unequivocal: teams were 23% more accurate than individuals. For year two, the researchers created special "superteams" of top forecasters, providing guidance on team functioning and online communication forums. Initially, many superforecasters felt intimidated by teammates with impressive credentials, and discussions were hampered by excessive politeness - what Marty Rosenthal called "dancing around" difficult topics.
Over time, team dynamics improved as members explicitly welcomed criticism and thanked others for constructive feedback. Without formal leadership structures, effective members like Rosenthal practiced "leading from behind" - setting examples of thorough analysis and organizing occasional conference calls. Face-to-face meetings strengthened team bonds, and members developed a strong sense of commitment.
Superteams divided workloads but paradoxically increased individual effort - members like Elaine Rich found teamwork "tons more work" but also "a rush." Teams excelled at information gathering, with members like Paul Theron contacting experts and sharing findings. The results were remarkable: superforecasters placed on teams became 50% more accurate, and superteams outperformed prediction markets by 15-30%, despite economists' expectations that markets would dominate.
Superteams succeeded by avoiding both groupthink and destructive conflict, creating "psychological safety" for challenging ideas respectfully. A team's actively open-minded thinking correlated with accuracy, but interestingly, team open-mindedness emerged from communication patterns rather than simply aggregating individual scores. The most successful teams fostered a culture of sharing, with "givers" like Marty Rosenthal, Doug Lorch, and Tim Minto contributing tools and analyses that benefited everyone - and these givers, far from being chumps, consistently ranked among the top individual forecasters.
Capítulo 10
The Leadership Challenge
Leaders must both make accurate forecasts and achieve goals through decisive action, creating an apparent contradiction. Effective leadership requires confidence, decisiveness, and vision - qualities that seem at odds with the superforecaster's uncertainty, deliberative thinking, and humility. How can leaders inspire confidence while seeing nothing as certain? How can they be decisive while engaging in slow, complex thinking?
The resolution to this leadership dilemma comes from 19th-century Prussian general Helmuth von Moltke, whose approach shaped modern military and corporate leadership. Moltke embraced uncertainty ("In war, everything is uncertain") and recognized that "no plan of operations extends with certainty beyond the first encounter with the enemy's main strength" - now simplified as "no plan survives contact with the enemy."
The Prussian military cultivated critical thinking through education that encouraged discussion and disagreement. Even junior officers could challenge generals' views. In extraordinary circumstances, disobedience was tolerated if justified - as when General Seydlitz refused King Frederick's order to attack prematurely, saying "Tell the King that after the battle my head is at his disposal, but meanwhile I will make use of it."
The key principle was Auftragstaktik (mission command) - pushing decision-making power down the hierarchy so those on the ground could respond quickly to battlefield surprises. Commanders told subordinates what goal to achieve but not how to achieve it, allowing officers to devise plans based on actual circumstances rather than headquarters' expectations.
While the Wehrmacht embodied Moltke's principles of independent thinking, America's post-WWI army initially rejected such approaches. When junior officer Eisenhower published an article advocating for tanks, he was threatened with court-martial for expressing ideas "incompatible with solid infantry doctrine."
Despite this culture, exceptional officers like Patton and Eisenhower valued individual initiative. Patton captured Auftragstaktik perfectly: "Never tell people how to do things. Tell them what to do, and they will surprise you with their ingenuity." As Supreme Commander, Eisenhower embodied Moltke's leadership - acknowledging uncertainty, maintaining a calm demeanor despite private anxieties, and encouraging open debate among officers.
Moltke's philosophy extends beyond military contexts into business leadership. William Coyne of 3M articulated this perfectly: "We let our people know what we want them to accomplish. But-and it is a very big 'but'-we do not tell them how to achieve those goals." Jeff Bezos's leadership principle "Have backbone; disagree and commit" similarly encourages respectful challenging of decisions while maintaining full commitment once decisions are made.
Capítulo 11
The Future of Forecasting
For months, the Scottish independence referendum seemed predictable, with "no" leading 57% to 43%. But two weeks before the September 2014 vote, polls shifted dramatically, giving "yes" the edge before another small shift put "no" slightly ahead with 9% still undecided. The outcome wasn't obvious until "no" won by a surprisingly wide 55.3% to 44.7% margin (a forecast superforecasters aced).
Political scientist Daniel Drezner broke from the typical pundit pattern of confident post-hoc analysis, instead asking the crucial question: "What does one do with data points like this to adjust one's worldview?" He recognized that without clear, scorable forecasts, we can't get the feedback needed for learning. Drezner committed to making clear predictions with confidence intervals going forward.
Forecasting tournaments may look like games with their scores and leaderboards, but the stakes are substantial. Good forecasting can determine business success or failure, policy effectiveness, and even peace or war. If US intelligence hadn't claimed certainty about Saddam Hussein's weapons of mass destruction, the disastrous Iraq invasion might have been avoided.
The strongest resistance to evidence-based forecasting comes from what Lenin called "kto-kogo" politics - the struggle for power where accuracy is secondary to advancing one's interests. When Nate Silver's forecasts favored Obama, Democrats praised him and Republicans reviled him; when he later predicted Republican Senate control, Democrats suddenly questioned his competence.
A century ago, physician Ernest Codman proposed his "End Result System" - tracking patient outcomes so hospitals could be evaluated on actual results rather than reputation. The medical establishment initially rejected him, even costing him his position at Massachusetts General Hospital. Yet his ideas eventually prevailed, becoming the foundation of evidence-based medicine.
This pattern of resistance followed by acceptance has repeated with evidence-based policy, charitable foundation evaluations, and sports analytics. All represent a broad shift from intuition and authority toward quantification and analysis. Even the intelligence community, which has strong incentives to avoid numerical precision in forecasting (to escape blame when wrong), funded the IARPA tournament - suggesting that history may favor evidence-based forecasting in the long run.
The post-2008 economic debate between Keynesians and Austerians exemplifies what's wrong with public discourse today. When Bloomberg reporters contacted signatories of a 2010 letter warning Fed Chairman Bernanke that his policies risked "currency debasement and inflation," they unanimously maintained they were right despite contrary evidence. Unlike the productive "adversarial collaboration" between psychologists Kahneman and Klein, these economic debates produce no learning.
My proposed solution: key figures could work together to identify precise, testable forecasts that meaningfully probe their disagreements - specifying exact inflation benchmarks and timeframes. Multiple questions would be needed, as complex theoretical disputes can't be resolved by single bets. The goal isn't gloating but learning - recognizing that reality is often more nuanced than either side initially believes. All we need to do is get serious about keeping score.