Chapitre 1
The Hidden Truths of Human Nature: What Google Searches Reveal About Us
Have you ever wondered what people search for when nobody's watching? While we carefully curate our social media personas and give socially acceptable answers on surveys, our Google searches tell a different story. In the privacy of our homes, we ask Google questions we wouldn't dare voice aloud-about our insecurities, prejudices, and darkest thoughts. This is the premise behind Seth Stephens-Davidowitz's groundbreaking work, which has been praised by Steven Pinker and featured in publications like The New York Times and The Economist. Drawing on his experience as a Harvard-educated data scientist and former Google analyst, Stephens-Davidowitz reveals how our digital footprints provide an unprecedented window into human nature. The book has become required reading in data science programs and boardrooms alike, offering insights that challenge conventional wisdom about everything from racism to sexuality. As Stephens-Davidowitz puts it, "The power of Big Data is that it gives us not just a new set of tools but a new way of understanding human behavior."
Chapitre 2
Digital Truth Serum: What We Hide from Others but Tell Google
Everybody lies. We lie to friends, bosses, spouses, doctors, and especially to surveys. This "social desirability bias" has plagued researchers for decades-people consistently misrepresent themselves to appear better than they are. A 1950 Denver study first exposed this phenomenon, showing significant gaps between self-reported behavior and official records on voting and charitable giving. Recent studies confirm this pattern continues today, with college graduates dramatically overstating their academic performance and donation history. For instance, surveys show that 40% of Americans claim to attend church weekly, while actual attendance records suggest the real figure is closer to 20%.
The digital age offers a solution: online searches function as "digital truth serum." Google searches meet all conditions for honesty-they're conducted online, alone, without an interviewer, and most importantly, with clear incentives to be truthful. When someone needs information about depression symptoms or racist jokes, they have genuine motivation to search honestly, unlike in surveys where there's no benefit to admitting socially undesirable thoughts or behaviors. This anonymity creates a unique window into human nature, revealing patterns that traditional research methods often miss.
This truth serum reveals surprising insights about sexuality. Traditional surveys suggest 2-3% of American men identify as gay, with significant variation between tolerant states like Rhode Island and intolerant states like Mississippi. However, pornography searches tell a different story-approximately 5% of male pornography searches nationwide are for gay content, with remarkably consistent rates across all states. Mississippi's 4.8% barely differs from Rhode Island's 5.2%, suggesting the true gay male population is around 5% nationwide, regardless of local attitudes. The consistency of these numbers across regions with vastly different cultural attitudes toward homosexuality suggests that sexual orientation is largely independent of social environment.
The data also exposes hidden struggles and internal conflicts. "Is my husband gay?" searches are particularly common in conservative states, with South Carolina, Louisiana, and Mississippi showing the highest rates. These searches spike by about 10% during major gay pride events and after related news stories. Craigslist casual encounter ads and "gay test" searches after viewing gay porn reveal complex patterns of sexual identity exploration. Beyond homosexuality, pornography searches challenge conventional wisdom about gender differences. Women's searches for violent content exceed men's by approximately 10%, contradicting traditional assumptions about gender and aggression. Google searches also reveal that men worry more about sexual rejection than women do, with twice as many complaints about boyfriends refusing sex than girlfriends. Men's searches about sexual performance anxiety peak late at night, suggesting these concerns are most acute in private moments.
The implications extend beyond sexuality to various aspects of human behavior. Search patterns reveal that people are more likely to admit prejudices, anxieties, and personal struggles online than in any other format. For instance, searches for racist jokes increase significantly after terrorist attacks, and searches about depression spike at 2 AM, providing insights into how external events and daily rhythms affect our most private thoughts and concerns.
Chapitre 3
Uncovering Hidden Prejudices: The Dark Side of Our Searches
Beyond sexual insecurities, Google search data reveals an uncomfortable truth about society's hidden prejudices - biases that people carefully conceal in public but express freely through anonymous searches. Americans frequently search questions like "Why are black people rude?" or "Why are Jews evil?", revealing distinct stereotypes that vary systematically by group. African Americans uniquely face "rude" and "angry" stereotypes, while nearly every minority group except Jews and Muslims encounters "stupid" stereotypes in searches. The "evil" label predominantly attaches to Jews, Muslims, and gay people, while Muslims alone consistently face terrorism-related stereotypes. These patterns persist across regions and demographics, suggesting deeply embedded societal prejudices.
These latent prejudices can rapidly transform into overt hatred during times of crisis or social tension. Following the 2015 San Bernardino shooting, searches for "kill Muslims" in California increased by over 400%, reaching a volume comparable to common casual searches like "martini recipe" or "weather forecast." The data revealed a particularly troubling pattern: when President Obama delivered a speech intended to calm Islamophobia and promote unity, it had the opposite effect. During and immediately after his address, searches calling Muslims "terrorists," "bad," and "evil" doubled in frequency, while "kill Muslims" searches tripled. Similar spikes occurred after other terror-related incidents, suggesting that attempts to reduce prejudice can sometimes intensify it.
The most disturbing findings center on searches containing the racial slur "nigger," which appears in approximately seven million American searches annually - equivalent to roughly 19,000 searches per day. Searches for "nigger jokes" vastly outnumber jokes targeting all other minorities combined, including Hispanic, Jewish, and Muslim groups. These searches demonstrate clear patterns, spiking dramatically whenever African Americans are prominent in news cycles - after natural disasters like Hurricane Katrina, during Obama's first presidential election campaign, and even during celebrations of Martin Luther King Jr. Day. The geographic distribution of these searches closely tracks historical patterns of racial animus.
This extensive search data fundamentally challenges the prevailing academic theory that modern racism persists primarily through implicit bias or unconscious associations. The prevalence and specificity of explicitly racist searches suggests that conscious, deliberate racism remains far more widespread than previously thought, merely hidden from public view. Statistical analysis reveals that these racist Google searches are actually better predictors of discrimination against Black Americans than traditional implicit association tests. Regions with higher rates of racist searches consistently show larger black-white wage gaps, increased incidents of police brutality, and stronger support for politically discriminatory policies.
Google searches also expose deeply ingrained gender biases, particularly in parenting. Parents search "Is my son gifted?" 2.5 times more frequently than they do for daughters, despite educational data showing girls are 9% more likely to be enrolled in gifted programs. Conversely, parents' searches about daughters disproportionately focus on physical appearance - there are twice as many searches about daughters being overweight (even though boys are statistically more likely to be overweight), and three times more searches about daughters being "ugly" or "unattractive." These biases show remarkable consistency across political affiliations, income levels, and geographic regions, and data since 2004 shows virtually no improvement in these gender-based search patterns.
Chapitre 4
Reimagining Data: Finding Value in Unexpected Places
The Big Data revolution isn't about collecting more data, but collecting the right data that provides unique insights. Google's remarkable success demonstrates this principle. While early search engines like AltaVista and Lycos counted word frequency (easily manipulated by keyword stuffing), Google's founders discovered a more valuable data source: the links between websites. By analyzing which sites others linked to when discussing topics, they effectively crowdsourced expertise across the internet. This link analysis proved incredibly predictive of useful information, propelling Google to dominance within two years of launch.
This approach extends beyond internet companies. In 2013, an unremarkable yearling horse (No. 85) with decent breeding and a concerning ankle scratch was set for auction at Fasig-Tipton. His owner, Egyptian beer magnate Ahmed Zayat, planned to sell him and buy other horses. But Jeff Seder's data-driven firm EQB made an unprecedented recommendation: don't sell this horse-he's potentially the best of the decade. Heeding this advice, Zayat secretly bought back his own horse for $300,000 (with 62 horses selling for more). Named American Pharoah, this horse would win the Triple Crown 18 months later.
Seder, a Harvard-educated eccentric who abandoned Wall Street for horse prediction, had spent decades measuring everything from nostril size to excrement volume, finding most measurements useless. His breakthrough came with portable ultrasound technology revealing that heart size-particularly left ventricle size-was the strongest predictor of racing success. American Pharoah's left ventricle measured in the 99.61st percentile, with all other organs proportionally large-a one-in-a-million combination.
Words have also become powerful data in the digital age. Google Ngrams reveals fascinating cultural shifts-like how Americans gradually shifted from saying "The United States are" to "The United States is," a change that occurred more gradually than historians previously thought, not immediately after the Civil War.
Even photos have transformed from physical objects into digital data, offering surprising insights when analyzed at scale. Computer scientists studying 949 digitized American high school yearbooks spanning 1905-2013 discovered a fascinating trend: Americans, especially women, gradually shifted from solemn expressions to beaming smiles. This wasn't because people became happier, but because Kodak's mid-century marketing campaign successfully associated photography with capturing happy moments.
Chapitre 5
Zooming In: Understanding Local Patterns and Individual Behavior
The author reveals striking geographic patterns in tax fraud behavior, highlighting how local knowledge networks influence financial decisions. In Miami, a remarkable 30% of self-employed individuals with one child reported exactly $9,000 in income-a suspiciously precise figure designed to maximize earned income tax credit-while Philadelphia showed only 2% exhibiting this pattern. This disparity wasn't rooted in regional differences in morality or ethics, but rather in information networks. People living near tax professionals or within communities where this knowledge circulated were significantly more likely to engage in this specific form of tax optimization. The author's research tracked individuals who relocated from low-fraud to high-fraud areas, documenting how they gradually adopted these practices, demonstrating that tax behaviors spread through social networks much like viral transmission.
Using an extensive dataset of 150,000 notable Americans from Wikipedia, the research uncovered profound geographic disparities in producing successful individuals. California-born baby boomers achieved Wikipedia notability at nearly four times the rate of their West Virginia counterparts. The success patterns clustered around two distinct types of locations: college towns and major metropolitan areas. College towns like Ann Arbor, Michigan, and Berkeley, California, showed exceptional ability to produce accomplished individuals across various fields. Particularly noteworthy was the role of historically black colleges-Tuskegee, Alabama, for instance, produced a disproportionate number of notable African American achievers despite its small size. Urban centers demonstrated specialized patterns: New York City excelled at producing journalists and media figures, Boston emerged as a hub for scientists and academics, while Los Angeles consistently generated successful entertainment industry professionals.
The concept of "doppelganger searches" represents a powerful analytical tool for prediction and pattern recognition. This approach revolutionized baseball analytics through Nate Silver's PECOTA system, which identified statistical twins to predict player trajectories. A compelling example emerged in 2009 when Boston Red Sox slugger David Ortiz faced a career crisis. While traditional metrics suggested his career was ending, doppelganger analysis identified similar players who had recovered from comparable slumps. The analysis proved prescient-Ortiz rebounded to have several more outstanding seasons, including leading the American League in slugging percentage in 2016.
The applications of doppelganger methodology extend far beyond sports. Tech giants have embraced similar principles: Amazon's product recommendations, Netflix's viewing suggestions, and Pandora's music matching all rely on sophisticated pattern matching algorithms. In medicine, this approach offers transformative potential. By analyzing vast databases of patient histories, doctors can identify "medical twins"-patients with similar conditions, treatments, and outcomes-to inform treatment decisions. This enables increasingly personalized medical care, moving beyond one-size-fits-all approaches to treatment protocols that consider individual patient characteristics and histories.
Chapitre 6
All the World's a Lab: Experiments in the Digital Age
On February 27, 2000, Google engineers had an idea that would revolutionize the internet: they conducted a simple experiment, randomly showing some users twenty search results instead of ten, then measuring user satisfaction. This seemingly modest test represented something profound-the digital world makes randomized controlled experiments (the gold standard for proving causality) incredibly cheap and fast compared to their resource-intensive offline counterparts.
This insight spread through Silicon Valley as "A/B testing," becoming the backbone of internet optimization. While traditional experiments might take months and thousands of dollars, digital A/B tests require just a line of code and automatic data collection. By 2011, Google was running 7,000 tests annually, and Facebook now conducts 1,000 daily-more experiments started in one day than the entire pharmaceutical industry begins in a year.
The power of A/B testing lies in overcoming our faulty intuition. When Obama's campaign tested different homepage combinations, the winning picture (Obama's family) and button text ("Learn More") generated 40% more sign-ups and $60 million in additional funding. News outlets like The Boston Globe test headlines constantly, finding that seemingly minor changes produce dramatic differences in clicks.
The dark side of this testing is its role in making the internet addictive. As "design ethicist" Tristan Harris notes, "There are a thousand people on the other side of the screen whose job it is to break down the self-regulation you have." Through relentless optimization, sites like Facebook and games like World of Warcraft become increasingly difficult to resist.
Natural experiments also help answer difficult causation questions. Using the 2012 AFC Championship game between the Patriots and Ravens, the author demonstrates how these real-world scenarios provide insights that controlled experiments cannot. When the Patriots won, Boston viewership of the Super Bowl increased by 60,000 people compared to Baltimore, creating a natural experiment for measuring ad effectiveness. Analysis showed Super Bowl ads generated a remarkable 2.8-to-1 return on investment for movie advertisements, suggesting companies are actually underpaying for these slots.
Chapitre 7
The Limits of Big Data: What It Cannot Do
The author explores the limitations of statistical inference through a compelling case study of Stuyvesant High School, one of New York City's most prestigious public schools. Using regression discontinuity analysis, economists compared students who scored just above and below the school's strict admission cutoff - sometimes separated by just a single test point. This natural experiment provided a unique opportunity to measure the true impact of elite education. The results revealed what researchers termed the "Elite Illusion" - students who barely made it into Stuyvesant performed no better on objective measures like AP exams, SAT scores, or college admissions than nearly identical students who just missed the cutoff and attended other schools. This finding challenged the deeply held belief that elite schools provide substantial academic advantages.
This insight was further reinforced by the landmark study from economists Stacy Dale and Alan Krueger, who tracked long-term outcomes of college students. They discovered that high-achieving students who were accepted to prestigious universities like Harvard but chose to attend less selective institutions like Penn State or state universities achieved similar career success and income levels. The research controlled for student ability and ambition by comparing students who were accepted to similar sets of colleges but made different choices. The results strongly suggested that individual characteristics like talent, drive, and work ethic were far more predictive of future success than institutional prestige.
The limitations of Big Data become even more apparent in financial markets, where sophisticated players deploy billions of dollars and advanced algorithms to exploit even minimal advantages. The author explains that the fundamental challenge lies in what statisticians call the "curse of dimensionality" - as the number of variables increases exponentially, the ability to find genuine patterns becomes increasingly difficult. For example, when analyzing a dataset with thousands of potential predictive factors against a limited number of historical observations, some correlations will inevitably appear significant purely by chance. This problem is particularly acute with modern Big Data sources, which can track countless variables from social media sentiment to satellite imagery of parking lots.
The author illustrates this through the example of the "lucky" coin phenomenon - if enough people flip coins to predict market movements, some will appear to have extraordinary predictive power simply due to random chance. When applied to complex financial markets, this same principle means that many apparent patterns or trading signals identified through data mining are likely to be spurious correlations rather than actionable insights. This becomes especially problematic as the sophistication and volume of data analysis increases, making it harder to distinguish genuine patterns from statistical noise.
Chapitre 8
The Ethical Challenges of Big Data
Language patterns in loan applications reveal surprising and sometimes unsettling predictors of repayment likelihood. Studies have found that applicants mentioning God, making emotional promises, expressing gratitude, or frequently referencing family members show a statistically higher probability of loan default. Conversely, applications containing specific repayment timelines, concrete financial planning details, and references to successfully fulfilled past commitments indicate greater reliability. These correlations persist across different demographic groups and loan types, from personal loans to business financing.
This discovery raises profound ethical questions about corporations leveraging our words and digital footprints to make life-altering decisions. Should companies deny loans based on statistically predictive but seemingly arbitrary linguistic markers? The implications extend far beyond finance into employment, housing, and insurance. In hiring practices, companies increasingly mine social media activity to evaluate candidates. Research from Cambridge University demonstrated that Facebook likes for classical composers like Mozart, scientific publications, or even specific foods like curly fries correlate with higher IQs, while preferences for certain brands like Harley-Davidson or particular lifestyle pages correlate with lower IQs. These correlations raise concerns about algorithmic discrimination and privacy rights.
The tragic case of Adriana Donato illustrates the darker implications of big data surveillance. Her murder by an ex-boyfriend followed weeks of his online searches about murder methods, disposal techniques, and alibi creation. This case exemplifies a broader ethical dilemma: should law enforcement use search data for preventive policing, and if so, where should we draw the line between public safety and privacy?
Google search patterns do show meaningful correlations with criminal activity across populations. Increases in suicide-related searches closely track actual suicide rates in specific regions, while spikes in Islamophobic searches correlate with subsequent anti-Muslim hate crimes. Using this aggregate data to inform resource allocation appears justified - cities experiencing surges in suicide-related searches might benefit from targeted mental health campaigns, while areas showing increased violent or discriminatory search patterns might warrant additional community protection measures.
However, using search data to target individuals presents both ethical and statistical challenges. The false positive rate is extremely high - while millions of people make suicide-related searches monthly, only a tiny fraction attempt suicide. Similarly, in 2015, there were approximately 12,000 searches for phrases like "kill Muslims" but 12 actual hate-crime murders of Muslims. This disparity highlights the danger of treating correlation as causation and the risk of creating pre-emptive surveillance systems that could infringe on civil liberties while generating overwhelming numbers of false alerts.
The challenge lies in balancing the potential benefits of predictive analytics with fundamental rights to privacy, presumption of innocence, and freedom from discrimination. As big data capabilities expand, society must establish clear ethical frameworks and legal boundaries for its use in decision-making that affects individual lives.
Chapitre 9
The Future of Data Science: A New Kind of Social Science
The future of behavioral science lies in what the author calls "science at scale"-taking simple methods and applying them hundreds of times using Big Data. From zooming in on health conditions to A/B testing educational software, we're discovering surprising results. For example, EDU STAR found that gamified math lessons actually performed worse than standard approaches for teaching fractions. Similarly, Jawbone discovered that a two-step nudge process could gain users an extra 23 minutes of sleep per night.
Psychology will increasingly adopt Silicon Valley's rapid testing methods, replacing months-long studies with undergraduates with thousands of digital tests performed in seconds. Text data will reveal profound insights about language development, idea spread, and humor. Children's online behavior, properly anonymized, could supplement traditional assessments of learning and development.
Karl Popper's critique extended beyond Freud to question whether any social science was truly scientific. When Popper compared physicists with social scientists, he found the former discovering deep truths while dismissing the latter as peddling gobbledygook. The Big Data revolution has changed this fundamental distinction. Today's economists and social scientists are answering clear questions with clear yes-or-no answers using mountains of honest data-the very definition of science.
Unlike physics with its simple, timeless laws, the social science revolution will likely come piecemeal through complex systems analysis. The datasets discussed in this book have barely been explored, with most academics still ignoring the digital data explosion. This will change as we apply techniques like John Snow's disease mapping at massive scale and utilize A/B testing for socially valuable outcomes like education and sleep improvement.
While digital truth serum reveals uncomfortable realities about human nature-from superficial judgments to hidden prejudices-this knowledge offers three significant benefits. First, it provides comfort in knowing we're not alone in our insecurities and embarrassing behaviors. Second, it alerts us to hidden suffering, enabling targeted interventions. Finally, and most powerfully, understanding these hidden truths can lead to solutions. As Stephens-Davidowitz concludes, "Google search data and other wellsprings of truth on the internet give us an unprecedented look into the darkest corners of the human psyche... We can use the data to fight the darkness. Collecting rich data on the world's problems is the first step toward fixing them."