Capítulo 4
The Power of What: How Correlation is Transforming Decision-Making
In 1997, Greg Linden joined Amazon and developed a revolutionary product recommendation system. While Amazon initially employed human reviewers to curate book selections, Linden realized an automated system comparing items with other items could be more effective than matching customers with similar customers. When both approaches were tested, the computer-generated recommendations dramatically outperformed human curation, leading to the dismissal of the reviewers. Today, Amazon's personalized recommendation engine generates one-third of the company's revenue and has transformed online retail. The crucial insight? Understanding WHY customers buy certain products together is unnecessary-knowing THAT they do is entirely sufficient.
This shift from causation to correlation represents the third fundamental transformation of Big Data. Correlations quantify statistical relationships between data points and enable predictions without understanding underlying mechanisms. Walmart employed this approach successfully in 2004 when they discovered that before hurricanes, customers bought not just flashlights but also Pop-Tarts. By placing both products strategically during storm warnings, they boosted sales without needing to understand the psychological reasons behind this shopping pattern.
Traditional hypothesis-driven approaches, where theories guide data collection, are giving way to data-driven approaches. Rather than laboriously testing individual hypotheses, advanced computational analysis can identify optimal correlations in enormous datasets. Google Flu Trends exemplifies this method, evaluating half a billion mathematical models to find the best predictive parameters for flu outbreaks.
The advantage is clear: we don't need to understand why people search for certain terms during flu outbreaks or why airline ticket prices fluctuate. By letting data speak for itself, we often obtain less biased, more accurate, and significantly faster results. Correlation analysis has become so ubiquitous that we barely notice it, though it forms the foundation of Big Data.
Companies like FICO leverage correlations to develop predictive models ranging from the classic credit score to the newer Medication Adherence Score, which predicts how reliably patients take their medications. These models incorporate seemingly irrelevant factors like residential stability or car ownership. Credit bureaus like Experian and Equifax have developed similar services that estimate income based on credit histories-much cheaper than traditional verification methods.
Perhaps most striking is insurance giant Aviva's approach, using credit and marketing data to predict health risks-an alternative to expensive blood and urine tests. For just $5 instead of $125 per applicant, the company can identify risks for hypertension, diabetes, or depression by analyzing hobbies, website visits, and television habits.
Target demonstrates this method's power most vividly: the retail chain can detect when a customer is pregnant based on purchasing patterns before her family knows. By analyzing shopping data, they identified about two dozen products that serve as early warning signals, including unscented lotion and certain supplements.
The shift from causation to correlation isn't without controversy. In 2008, Wired editor Chris Anderson provocatively claimed the "end of theory"-that the scientific method would be replaced by statistical correlation analyses without theoretical foundation. He argued that with enough data, "the numbers speak for themselves" and "correlation is enough."
This claim is overstated. Even Big Data relies on theories-statistical, mathematical, and computational. Our theories shape data selection, analysis methods, and result interpretation. Google used search queries for flu predictions, not hair lengths-a theory-guided decision.
Nevertheless, Anderson deserves credit for raising important questions. Big Data may not signal the "end of theory," but it fundamentally changes our worldview. This transformation challenges many institutions, yet the immense benefits make Big Data inevitable. The deeper reason for this revolution lies not just in new digital tools but in the vast data availability-which itself results from the increasing transformation of reality into data.
Capítulo 5
The Datafication Revolution: Turning the World into Information
When Matthew Fontaine Maury, a disabled naval officer, took charge of the U.S. Navy's Depot of Charts and Instruments in the 19th century, he discovered a treasure trove of neglected logbooks containing information on winds, waves, and weather. Recognizing their potential value, he systematically organized this data, dividing the Atlantic into five-degree quadrants and recording temperatures, wind speeds, and wave conditions by month.
To gather more data, Maury designed standardized log forms and created an early "social network" with merchant ships, providing them with his tables in exchange for their logbooks. The results were remarkable: his tables shortened overseas voyages by an average of one-third. "Before I had your work in hand, I sailed the seas with my eyes closed," one captain wrote to him. By 1855, Maury's "Physical Geography of the Sea" had processed 1.2 million data points.
Maury's work exemplifies datafication-the transformation of phenomena into analyzable data. This process distinguishes primitive from advanced societies. Simple counting and measuring began millennia ago in early civilizations. The invention of writing in Mesopotamia enabled precise recording of transactions and made human actions repeatable-buildings could be reconstructed from records, economic transactions became traceable.
Over centuries, measurement expanded to more domains. The introduction of Arabic numerals in Europe (originally from India via the Arab world) revolutionized mathematics. Unlike Roman numerals, they allowed efficient calculation. Double-entry bookkeeping, emerging in 14th-century Italy and popularized by Luca Pacioli's 1494 textbook, was another datafication milestone, enabling standardized records and making profits and losses immediately visible.
The Medici family used these techniques to become Europe's most influential bankers. In the 19th century, measurement techniques became increasingly precise, aiming to understand nature through quantification. The invention of computers finally made datafication vastly more efficient-digitization became datafication's turbocharger, though they're not synonymous.
Google's book project illustrates the difference between digitization (book pages as images) and datafication (text as processable data). Through optical character recognition, Google transformed scanned pages into searchable, analyzable data, enabling new applications: text search, historical word analysis, and the "Ngram Viewer" that graphically displays word frequency over time. By 2012, Google had scanned 20 million books (15% of humanity's written heritage). This gave rise to "culturomics"-studying cultural developments through text analysis. Harvard researchers discovered that less than half of all English words appear in dictionaries. Unlike Amazon, which primarily views e-books as reading content, Google uses datafied texts for diverse analyses like improving translation services.
The datafication of location required three prerequisites: a system for measuring the Earth, a standardized recording method, and data collection tools. This journey began with Eratosthenes' grid lines (200 BCE), continued with Ptolemy's Geography and Mercator's conformal projection (1570), and was formalized in 1884 with the Greenwich Prime Meridian standardization. The breakthrough came in 1978 with the GPS satellite system, enabling precise positioning. As GPS module costs fell to about one dollar, continuous location tracking became ubiquitous. Companies like UPS now use "geo-loco data" for route optimization, saving 50 million kilometers, 12 million liters of fuel, and 30,000 tons of CO2 in 2011 alone. Insurance companies also use location data for individual risk calculation, fundamentally changing the traditional insurance concept.
Social media forms the backbone of personal relationship datafication. Facebook has datafied relationships with its "Social Graph," Twitter captures fleeting thoughts and feelings, while LinkedIn transforms professional experiences into usable data. With a billion users and over 100 billion friendships, Facebook's database represents about 10% of the world's population. The applications are enormous: startups already use the Social Graph for credit checks, based on the assumption that people with similar characteristics preferentially befriend each other.
Twitter data (400 million daily tweets from 140 million monthly users) is marketed by companies like DataSift and Gnip. Hedge funds like Derwent Capital and MarketPsych use "sentiment analysis" for investment decisions, while Thomson Reuters offers 18,864 sentiment indices from 119 countries. Researchers have discovered cross-cultural mood patterns by analyzing 509 million tweets and can even predict behaviors like vaccination willingness.
Datafication increasingly extends to all aspects of life and will fundamentally change our society-similar to historical infrastructure projects. While aqueducts, printing presses, and the internet enabled the flow of water or knowledge, datafication represents a fundamental enrichment of human understanding. With Big Data, we no longer see the world as a sequence of events but as a universe essentially consisting of information.
Physics has taught us for over a century that information forms the basis of everything. Through datafication, we can more comprehensively measure, calculate, and act upon all aspects of our existence. Future generations will likely develop "Big Data consciousness"-the natural assumption that everything has a quantitative component and that data is an indispensable source of learning. The effects of datafication may even surpass those of the printing press and internet by giving us the means to understand the world as quantifiable and analyzable.
Capítulo 6
Unlocking Hidden Value: The Economics of Data Reuse
The true value of data resembles an iceberg-only a small portion is initially visible. As Farecast, Google, and Maury demonstrated, the real value often lies in secondary uses far beyond the original purpose.
Data behaves like potential energy in physics-stored and dormant until applied to a new purpose. The "option value" of data is the sum of all possible uses. In the Big Data era, data isn't a one-time resource but a "magical diamond mine" that yields returns repeatedly.
There are three main ways to tap this option value: reuse, recombination, and double-purposing. In reuse, data serves new purposes-like Google's search queries for economic forecasts or Telefonica selling anonymized user location data. Many companies sit on unused "data troves" whose value they don't recognize.
Recombination combines different datasets, as in the Danish cancer study that linked mobile data with cancer registries. Double-purposing collects data to serve multiple purposes from the start-like surveillance cameras that not only prevent theft but also analyze customer flows.
However, data loses value over time. Outdated data can even diminish newer data's value, as with Amazon's product recommendations. The challenge lies in recognizing which data remains relevant and which should be discarded.
Data reuse creates enormous added value, as internet search engines demonstrate. What appears as a one-time interaction becomes valuable market research: Hitwise analyzes search terms for businesses, Google cooperates with banks for economic forecasts, and the Bank of England uses property-related searches for market analysis.
Some companies recognize this potential too late: AOL didn't understand that Amazon was collecting valuable user data through their partnership. Similarly, Nuance provided Google with speech recognition software without contractually securing rights to the voice recordings.
Traditional companies can also benefit from their data treasures. A logistics company developed a lucrative economic forecasting service from its delivery data, while SWIFT generates GDP forecasts from bank transfer data. Mobile operators like Telefonica have recognized that their technically necessary location data can be sold as anonymized aggregates to retailers.
The real value of data often emerges only through innovative combination of different sources. A remarkable study on the connection between mobile phone use and cancer risk demonstrates this principle impressively. Danish researchers combined mobile data of all participants since 1987 with the national cancer registry and socioeconomic data.
This method overcame typical weaknesses of previous studies: instead of small samples, 358,403 mobile users and 10,729 cancer patients were included-almost following the "N = all" principle. The data was high-quality, independently collected, and enabled analysis of 3.8 million person-years of mobile usage.
With Big Data, the sum is more valuable than its individual parts. While simple "mashups" like Zillow's property price maps are already familiar, the Danish study shows the full potential of data combination-even though it ultimately found no correlation between mobile phone use and cancer risk.
Data reuse becomes significantly easier when collection is designed for multiple purposes from the outset. Retailers position surveillance cameras to not only prevent theft but also analyze customer flows-transforming them from pure security measures to revenue-enhancing investments.
Google exemplifies this strategy with its Street View vehicles, simultaneously collecting photos, GPS data, map data, and WiFi information. This "extensible" data was conceived from the beginning for multiple uses, such as mapping services and self-driving cars.
Since additional data streams often add minimal cost, it makes economic sense to collect as much data as possible and look for "twofers"-cases where data can serve multiple purposes. This maximizes the option value of the collected information.
Capítulo 7
The Rise of Data Science: New Players in the Information Economy
The Big Data ecosystem has developed three types of companies, distinguished by their value creation:
1. Data owners: Companies with access to valuable data but not necessarily the skills to extract value (like Twitter, which employs other firms to license its data).
2. Capability owners: Consulting firms, technology providers, and data analysts with expertise but without their own data or innovative ideas (like Teradata, which conducts analyses for Walmart).
3. Idea generators: Companies whose success is based on unique ideas for creating value from data, like Pete Warden's Jetpac, which makes travel suggestions based on user photos.
While capabilities and data ownership were initially the focus, "data scientist" has emerged as a new profession-combining statistician, software developer, infographic designer, and storyteller. The McKinsey Global Institute predicts a continued shortage of these specialists.
The Big Data value chain begins with data holders who control access to information. Some use it themselves, others license it to third parties. The most successful companies like Google combine all three elements of the value chain: they own data, have innovative ideas, and possess the necessary implementation skills. Google collects typing errors in search queries, develops spell-checking from them, and makes parts of its data available to others through APIs.
Amazon follows a different strategy-it first developed the idea (like its book recommendation system) before collecting the necessary data. While Google considers secondary uses when collecting data (like GPS data for mapping services and self-driving cars), Amazon focuses on primary benefits and treats secondary uses as a bonus. Amazon's Kindle records which book pages are frequently annotated but doesn't share this valuable information with authors or publishers.
Big Data is also changing traditional business relationships. A European car manufacturer used sensor data from its vehicles to identify defects in supplier parts, developed an improved solution, patented it, and sold the patent to the supplier-instead of simply sharing the information.
Although innovative ideas and capabilities appear most valuable in the early phase of the Big Data era, the greatest long-term value will concentrate in the data itself. Increasingly, "data intermediaries" are emerging that combine information from various sources to create new value.
Companies are experimenting with different organizational forms for Big Data. Inrix was deliberately designed as a neutral data intermediary to encourage competing companies to collaborate. Similarly, UPS sold its internal data analysis department as Roadnet Technologies, which now conducts traffic analyses for various companies-with the advantage that competitors provide their data, which they would have refused to a UPS subsidiary.
The most valuable component is ultimately the data itself, not the expertise or attitude. This is also shown by company acquisitions: While Microsoft bought Farecast for $110 million in 2006, Google paid $700 million for ITA Software, Farecast's data supplier, in 2008.
Capítulo 8
The End of Expertise: How Data is Replacing Human Judgment
Big Data is increasingly replacing human judgment with data-driven decisions. As shown in the film "Moneyball," subjective expert opinions are being supplanted by objective data analyses. The film illustrates how baseball talent scouts previously made decisions about million-dollar player contracts based on gut feeling, using superficial criteria like physique or even a player's girlfriend's appearance.
Billy Beane revolutionized this approach as manager of the Oakland A's by replacing traditional evaluation methods with new, mathematically grounded metrics. Against considerable resistance, he introduced "sabermetrics" and transformed a mediocre team into a championship winner.
This transformation is occurring across numerous industries: Online media like Huffington Post and Forbes let data rather than editors determine content. Coursera uses usage data to improve course materials. Amazon replaced its book reviewers with automated recommendation systems.
The most successful Big Data pioneers are often outsiders-specialists in data analysis, AI, or statistics who apply their knowledge to unfamiliar industries. In Kaggle competitions, industry outsiders frequently develop the best algorithms: A British physicist created successful prediction models for insurance claims, an accountant from Singapore won a competition for predicting biological responses.
This development fundamentally changes which skills are necessary for professional success. Instead of specialization and depth, breadth and data understanding increasingly count. Mathematics, statistics, and programming are becoming basic competencies like reading and writing were previously.
The video game industry exemplifies how Big Data is displacing traditional expertise. Previously, designers created games based on creative intuition and hoped for success. Today, Zynga, maker of FarmVille and other online games, continuously analyzes user data to optimize games.
Zynga doesn't just adapt games generally but creates hundreds of different versions for individual players. When data analyses showed that FishVille players bought transparent fish six times more often than other types, Zynga expanded the offering accordingly. In Mafia Wars, they discovered the preference for gold-plated weapons and white tigers as pets-insights no designer would have gained through intuition alone.
"We're an analytics company masquerading as a game maker. Everything depends on the numbers," explained Ken Rudin, former chief analyst at Zynga, who later moved to Facebook.
The shift to data-based decisions is profound. The-Numbers.com demonstrates this change in the film industry by analyzing complex correlations from 30 million data points. The analyses provide concrete recommendations-like hiring an Oscar-nominated actor for $5 million or reducing a sailing film's budget from twelve to eight million.
Studies by MIT professor Erik Brynjolfsson show that companies with strongly data-driven decision-making demonstrate up to 6 percent higher productivity than those with intuitive decision processes-a significant competitive advantage, though one that might diminish with the increasing spread of Big Data methods.
Capítulo 9
The Dark Side: Privacy, Prediction, and the Dictatorship of Data
For nearly forty years, the Stasi monitored millions of East German citizens with 100,000 employees, collecting 39 million index cards and 100 shelf-kilometers of files. Today, more data is collected about each of us than ever before-through credit card payments, mobile phones, surveillance cameras, and the internet. Specialist companies like Equifax, Experian, and Acxiom collect personal data on hundreds of millions of people worldwide.
Big Data threatens privacy more severely than the internet and risks punishing people for predictions before they've even acted-potentially ending fairness, justice, and free will. A third danger is the "dictatorship of data," when information becomes fetishized and can be misused.
History shows how data has been misused for "bloody purposes": The U.S. Census Bureau provided addresses of Japanese-American citizens for their internment in 1943, Dutch population registers were used by Nazis to identify Jewish citizens, and IBM punch card systems organized the Holocaust. Today's technical possibilities far exceed those of the Stasi, and police departments already use algorithmic models for patrol planning.
The risks of Big Data grow with the data volumes themselves. While the Stasi once had to collect data with great effort, today much of this information is routinely captured by mobile providers. Police departments already use algorithmic models for their patrol planning, indicating the development direction.
Not all Big Data sets contain personal information-sensor readings in refineries or factories are harmless in this respect. But many modern data are personal, and economic incentives drive companies to collect more and store longer. Even seemingly anonymous data can often be traced back or allow inferences about confidential details.
Smart meters, for instance, record electricity consumption at short intervals and can identify the typical "load curves" of various electrical devices-from PCs to cannabis growing lamps-revealing residents' daily routines, health, or illegal activities.
In the Big Data era, traditional data protection methods fail: Consent declarations don't work when the future use of data is still unknown. Opt-out procedures leave traces, and anonymization becomes increasingly ineffective due to the abundance of data, as cases at AOL and Netflix showed, where "anonymized" users could easily be identified.
Surveillance is now easier, cheaper, and more comprehensive than in East German times. Although companies don't have state power, they hoard enormous amounts of personal data and use them for barely imaginable purposes. Meanwhile, government agencies like the NSA collect billions of communication records daily and store them in gigantic data centers-not to constantly monitor everyone, but to access them immediately when needed.
Like the opening scene of "Minority Report," where people are arrested for crimes before committing them, a disturbing future vision of Big Data use emerges: People could be declared guilty based on behavioral predictions without ever having committed a crime.
Early approaches to this development already exist. In more than half of U.S. states, parole boards use data-based behavioral predictions for their decisions. "Predictive policing" in cities like Los Angeles and Richmond monitors certain streets, groups, or even individuals more intensively just because an algorithm has identified them as more susceptible to crime.
In Memphis, the Blue CRUSH program provides police with precise information on where and when to concentrate their forces. Since its introduction in 2006, the number of serious offenses has reportedly decreased by a quarter. Richmond correlates crime statistics with additional data such as payroll dates or major events, refining police assumptions-showing, for instance, that the increase in violent crime after gun shows occurs only two weeks later.
While traditional profiling is based on group characteristics and can lead to discrimination, Big Data promises individualized predictions. This initially appears as progress but becomes problematic when people should be punished for acts not yet committed. This contradicts the basic principle of justice: One must have done something before being held accountable.
Perfect predictions would negate free will. In reality, however, Big Data only provides probabilities. Professor Richard Berk's method for predicting homicides in parole cases achieves about 75 percent accuracy-meaning a quarter of decisions would be wrong. Moreover, with preventive intervention, it's impossible to verify whether the prediction would have actually come true.
Using Big Data to assign blame for predicted actions undermines the presumption of innocence and our ability to make moral decisions. This danger extends far beyond state prosecution to all areas of life-from termination decisions to medical treatments. Big Data is useful for risk assessments but unsuitable for causal blame assignments since it's based on correlations rather than causality.
Capítulo 10
Controlling the Future: New Rules for a Data-Driven World
When information systems evolve, societal rules must adapt. The Gutenberg press exemplifies this: before its invention, knowledge was limited to monasteries and small university collections. Its spread across Europe revolutionized information flow, leading to new control mechanisms like censorship, printing patents, and copyright laws.
In the Big Data era, we face a similar shift, but with years instead of centuries to respond. The traditional data protection principle of individual consent no longer suffices, as data's value often emerges through unexpected reuse. Instead, responsibility should shift to data users, who would conduct impact assessments for new purposes under legislative guidelines.
This approach allows companies to repurpose data without repeated consent but holds them legally liable for inadequate assessment. For example, a car manufacturer using seat recognition for theft protection could extend it to detect driving incapacity after proper risk evaluation.
To protect free will in the Big Data era, we must maintain that people are judged by actions, not predictions. While Big Data analysis can identify potential cases, evidence must be gathered traditionally. For significant private sector decisions, protective measures should include algorithm disclosure, independent certification, and opportunities to challenge predictions.
As Big Data systems grow more complex, transparency becomes crucial. Systems like Google Translate use billions of data points, making their decision-making process increasingly opaque. To maintain accountability, we need "algorithmists" - experts in computer science, mathematics, and statistics who evaluate Big Data analyses.
These algorithmists would function similarly to auditors, sworn to confidentiality and independence. They would examine data selection, analysis tools, and result interpretation. They could serve as:
• External examiners verifying predictions' accuracy
• Court experts in complex cases
• Consultants for individuals affected by Big Data decisions
• Internal monitors within companies, similar to data protection officers
This professional oversight would help balance innovation with accountability, ensuring Big Data serves society while protecting individual rights. Like data protection officers in Germany, algorithmists would maintain professional ethics while serving both their employers and their obligation to independence.
Capítulo 11
The Human Element: Finding Balance in a Data-Driven World
In an increasingly data-driven world, the crucial question remains: What role do intuition, faith, uncertainty, and originality still play? Big Data teaches us that we can act successfully without understanding everything. Flowers and his team in New York save lives without being omniscient experts.
The Big Data world is not a cold, algorithm-controlled place. People with their weaknesses, misunderstandings, and errors play an essential role, as these characteristics are inseparably linked to creativity, instinct, and genius. In a world of data-based decisions, the human element-intuition, risk-taking, and serendipitous discoveries-becomes the decisive differentiating factor.
For innovation, Big Data is a powerful tool, but the actual spark of inventive spirit lies beyond the data. As Henry Ford noted: Had he asked his customers about their wishes, they would have only requested "a faster horse." We must especially nurture our most human qualities-creativity, intuition, and intellectual ambition-as they are the true source of progress.
Big Data always remains incomplete. Even CERN collects less than 0.1 percent of the data generated in its experiments. Our insights will always be limited by the constraints of our measuring instruments. What we call "Big Data" today will soon seem as outdated as the four kilobytes of RAM in the Apollo 11 onboard computer. Our data collection will always be just a simulation of reality-like the shadows in Plato's cave allegory. This doesn't make Big Data worthless but reminds us to use this tool with humility and humanity.
As we navigate this new era, we must remember that data is not an end in itself but a means to better understand and improve our world. The true power of Big Data lies not in replacing human judgment but in enhancing it-providing insights that allow us to make better decisions while preserving the uniquely human qualities that give those decisions meaning and purpose.