Глава 1
The Data Revolution: How Big Data is Transforming Business and Society
Have you ever wondered how Target knew a teenage girl was pregnant before her own father did? Or why your car insurance company suddenly wants to track your driving habits? Welcome to the world of Big Data-a revolution that's transforming how businesses operate, governments function, and individuals live. Phil Simon's "Too Big to Ignore" arrived at a watershed moment when Big Data was transitioning from tech industry buzzword to mainstream business imperative. The book has become required reading in MBA programs nationwide and has influenced how organizations approach their data strategies. Even celebrities like Ashton Kutcher, an active tech investor, have praised the book's accessible approach to a complex topic. As data continues to grow exponentially-with humans now generating more information in two days than was created from the dawn of civilization through 2003-understanding Big Data isn't just advantageous; it's essential for survival in today's information economy.
Глава 2
From Structured Data to the Data Deluge: The Evolution of Business Information
The data landscape has undergone a dramatic transformation in recent decades. Before office computers became standard, companies relied on manual processes for everything from payroll to inventory management. The first major leap came with VisiCalc in the mid-1970s, the pioneering spreadsheet program that revolutionized business calculations. By the mid-1980s, enterprise systems like MRP and ERP emerged, processing structured data-orderly information about customers, employees, and products organized in relational tables.
Everything changed around 2005 with the emergence of Web 2.0-the social web. Suddenly, the volume, variety, and velocity of data exploded exponentially. While structured data continued growing impressively (Walmart's databases expanded from 100 terabytes to 2.5 petabytes between 2005-2011), unstructured data grew 10-50 times faster, now accounting for over 80% of all organizational information.
Unlike structured data that fits neatly into tables, unstructured data is messy, text-laden, and nonrelational. It includes everything from social media posts to customer reviews to video content. Semi-structured data occupies a middle ground, with characteristics of both types-think XML documents, emails, and EDI standards for computer-to-computer exchange.
Metadata-data about data-has become increasingly valuable. When you upload photos to sites like Flickr, you're not just sharing images but creating rich metadata including timestamps, file sizes, and sometimes GPS coordinates. This metadata enables powerful search capabilities and connections between seemingly unrelated content.
Despite this data revolution, most organizations continue struggling even with basic data management. They face poor data quality, lack master data management, and operate without proper data governance. This creates a pyramid structure where only a small percentage of companies effectively manage information while the majority waste time answering what should be simple business questions about headcount, sales figures, or inventory levels.
The fundamental disconnect many organizations face is trying to address new forms of data with tools designed for traditional data types. Fewer than half of companies currently collect and analyze social media data, and standard reports and business intelligence applications weren't built to handle unstructured information. The opportunity lies in embracing new tools specifically designed for Big Data's unique characteristics.
Глава 3
Understanding Big Data: Beyond Volume, Variety, and Velocity
Like Justice Potter Stewart's famous "I know it when I see it" threshold for obscenity, Big Data lacks a perfect definition despite its widespread usage. Douglas Laney's 2001 framework of three dimensions-volume (increasing amount), variety (expanding types), and velocity (accelerating speed)-remains foundational, though some have added dimensions like veracity and variability.
Rather than seeking a perfect definition, understanding Big Data's characteristics proves more useful. First, Big Data isn't theoretical-it's already transforming industries from politics to winemaking. The 2012 Obama campaign demonstrated its power through sophisticated analytics that helped raise over $1 billion and secure victory. Companies like Gallo launch new wine varietals based on social network analysis, while UPS optimizes delivery routes through sensor data.
Today's media landscape is incredibly fragmented. Unlike when 125 million Americans watched the M*A*S*H finale together, we now consume increasingly diverse and niche content across platforms. We generate and interact with more data than ever, but each person only engages with a tiny fraction of available content-over 70% of tweets are ignored, and despite having thousands of social connections, Dunbar's Number limits meaningful relationships to about 150 people.
Big Data isn't a magical solution to organizational problems. It won't fix broken corporate cultures or incompetent management. A Forbes survey found over 60% of knowledge workers at large enterprises lack processes and skills to use information effectively for decision making. Organizations that embrace Big Data may gain competitive advantages, but success isn't guaranteed.
Effectively managing internal structured data ("Small Data") remains essential. Companies that excel at managing their Small Data will benefit more from Big Data initiatives. While Small Data serves a descriptive purpose (showing what's happening now), organizations gain the most value when they can integrate it with Big Data for predictive capabilities.
Big Data complements but doesn't replace traditional systems. Like a golf driver that can hit balls 300 yards but isn't suitable for putting, Big Data excels at certain tasks but can't perform essential organizational functions like generating P&L statements. Organizations adept with both Small and Big Data have the greatest advantage.
While Big Data improves forecasting capabilities beyond what's possible with historical data alone, it cannot eliminate uncertainty entirely. Netflix exemplifies both the power and peril of Big Data. The company built a beloved recommendation engine using collaborative filtering, tracking billions of customer interactions. Yet when Netflix made the ill-conceived Qwikster decision in 2011, Big Data accelerated its downfall as customer outrage spread through social media, causing the company to lose 800,000 customers and half its market value.
Unlike traditional relational databases that contain many rows (records) and relatively few columns (fields), Big Data operates differently, using columnar databases that are more wide than long, making them better suited for handling unstructured information beyond the capabilities of traditional SQL.
Perhaps the most clarifying approach to understanding Big Data is defining it by what it isn't. According to The Register, "Big Data is any data that doesn't fit well into tables and that generally responds poorly to manipulation by Structured Query Language (SQL)." The fundamental distinction is that Big Data doesn't fit neatly into traditional database structures.
Глава 4
The Analytical Arsenal: Techniques That Power Big Data Insights
At its core, Big Data enables organizations to accomplish three critical objectives: better understand the past (what happened and why), better comprehend the present (what is happening and why), and better predict the future (what will happen and why). While Big Data cannot change the past or completely alter the present, it significantly enhances our ability to make accurate predictions by reducing uncertainty.
Statistical techniques remain foundational in the Big Data era. Regression analysis allows organizations to understand relationships between variables like product sales and marketing spend across massive datasets. A/B testing compares different content variations against a control, offering a scientific approach to decision-making. Companies like CapitalOne run thousands of tests annually to optimize offers and maximize profits. The beauty of A/B testing is that it generates empirical data that can't be argued with, removing subjective opinions from the equation.
Data visualization has evolved far beyond basic Excel charts into sophisticated, interactive tools that help organizations understand massive datasets. Heat maps provide visual representations where values are depicted by color intensity, while time series analysis examines data points ordered by time to identify trends beneath seasonal variations. For businesses like retail stores, this helps distinguish between normal seasonal patterns (like holiday sales spikes) and actual business changes.
Modern technology has elevated automation to unprecedented levels, enabling Big Data capabilities. Machine learning gives computers the ability to discover patterns without explicit programming. Companies like Twitter, Google, and Facebook use machine learning to process enormous amounts of data-from filtering millions of comments to detecting credit card fraud. Advances in sensors and nanotechnology dramatically expand data capture capabilities. A Boeing jet engine produces 10 terabytes of operational data every 30 minutes, with a four-engine jumbo jet generating 640 terabytes during a single Atlantic crossing.
Semantics technologies translate unstructured text-based data into meaningful information. Natural Language Processing (NLP) produces readable summaries from text chunks, with applications ranging from social media translation to healthcare. In medicine, NLP can decipher doctors' notes, unlocking the estimated 80% of clinical documentation that exists as unstructured "text blobs." Some NLP applications have achieved over 90% accuracy in identifying diseases from doctors' text descriptions before any lab testing.
Text analytics helps organizations make sense of unstructured data from sources like customer reviews, emails, and call center transcripts. It transforms information retrieval into information access by mining retrieved material for structure, entities, topics, and relationships. Sentiment analysis examines text to characterize its tonality, allowing organizations to understand not only how people feel but also the strength of their feelings and underlying reasons.
Predictive analytics represents the Holy Grail of Big Data-forecasting probabilities by combining multiple variables into statistical models that improve as additional data becomes available. Google's search prediction capabilities exemplify this, suggesting corrections when users make typing errors and predicting flu outbreaks through geographic search pattern analysis.
Two fundamental principles govern Big Data analytics. First, the Law of Large Numbers states that as sample size increases, the average approaches the population mean-essentially, predictions become more reliable with larger datasets. Second, the Law of Diminishing Marginal Utility explains how each additional data point provides less value than the previous one. However, today's low storage costs mean organizations can now capture even marginally valuable data that might reveal emerging trends.
Collaborative filtering combines large number statistics, powerful technology, and crowdsourcing to generate remarkably accurate product recommendations. Companies like Amazon and Netflix analyze user ratings and behaviors across millions of customers to identify patterns and make personalized suggestions. The beauty of collaborative filtering is its resilience to outliers-even when users have unusual preferences, large sample sizes ensure the overall recommendations remain accurate.
Глава 5
The Technology Backbone: Solutions That Make Big Data Possible
Traditional database systems like RDBMSs can't efficiently handle Big Data's volume, variety, and velocity. Organizations need specialized tools operating at a completely different scale to implement techniques like sentiment analysis or predictive modeling.
Hadoop stands as the de facto Big Data platform-an open-source collection of projects that distributes and processes vast amounts of data across multiple servers. Used by giants like Yahoo!, Facebook, LinkedIn, and Twitter, Hadoop's popularity stems from its ability to handle diverse data types, scale horizontally, maintain high fault tolerance, and remain extremely flexible.
Its foundation rests on MapReduce, which breaks Big Data problems into manageable subproblems, distributes them to processing nodes, and reaggregates them into digestible datasets. The Hadoop Distributed File System (HDFS) works with components like HBase to store and provide access to enormous datasets spanning billions of rows and millions of columns.
Hadoop's open-source nature has spawned a vibrant ecosystem of complementary projects and extensions. Companies like Cloudera, Hortonworks, and MapR have built commercial offerings around Hadoop, providing enterprise-grade support, services, and training. Innovative startups like Hadapt and Platfora are building on Hadoop's foundation to solve specific Big Data challenges, while established enterprise vendors like Oracle, IBM, and Microsoft have developed their own Hadoop-compatible products.
Despite its power, Hadoop has notable limitations. It doesn't provide real-time information (though HBase and Impala help close this gap), requires significant programming expertise, and may increase security risks through data consolidation. The lack of formal industry standards remains a concern, though Hadoop continues evolving rapidly.
As Big Data doesn't play well with traditional relational databases and SQL, organizations need alternative storage solutions. The NoSQL movement offers alternatives across four main types: Key-Value Stores (like Redis), Column Family Stores (like Cassandra), Document Databases (like MongoDB), and Graph Databases (like Neo4J). Organizations often use both SQL and NoSQL databases concurrently for different purposes.
NewSQL represents a new generation of relational databases attempting to combine NoSQL's speed and scalability with traditional SQL systems' proven capabilities. Companies like VoltDB are creating SQL 2.0 architectures that address the limitations of traditional RDBMSs while maintaining SQL's familiarity.
Columnar databases address fundamental limitations of row-based systems when handling massive datasets. While traditional RDBMSs examine every field in every row, columnar databases only access the specific columns needed, dramatically reducing I/O bottlenecks. This approach offers 7-8x better data compression and vastly superior performance for analytical workloads.
For organizations wanting to experiment with Big Data before making major commitments, cloud-based analytics services like 1010data can process trillions of records quickly. Kaggle represents a fascinating hybrid business model combining crowdfunding, crowdsourcing, and gamification to match organizations possessing data but lacking analytical expertise with a community of over 40,000 data scientists worldwide.
Глава 6
Real-World Impact: How Organizations Are Transforming Through Big Data
While theoretical discussions about Big Data abound, examining real-world applications reveals its true transformative potential. Three diverse organizations demonstrate how Big Data delivers concrete results regardless of company size or industry.
Quantcast, founded in 2006, analyzes over 300 billion observations of media consumption monthly to connect advertisers with their ideal customers. Recognizing early that traditional data approaches wouldn't suffice, they built their own distributed file system (QFS) that outperforms Hadoop while using 50% less disk space. Their customer success stories include a national auto parts retailer achieving 200% ROI, a wireless company increasing conversion rates by 76%, and a hotel chain doubling bookings. Despite being relatively small with 250 employees, Quantcast proves that embracing Big Data from inception yields significant competitive advantages.
Explorys tackles the unsustainable $3 trillion U.S. healthcare system where an estimated $1.2 trillion is wasted annually. Their DataGrid platform integrates clinical, financial, and operational healthcare data using Hadoop-based technology to improve care delivery while reducing costs. With nearly 100 employees, Explorys serves major healthcare networks like Cleveland Clinic and MedStar, providing subsecond search across patient populations and enabling proactive care coordination.
Their practical applications include helping hospitals understand why patients use emergency rooms for non-emergency care by analyzing demographics, medical histories, and neighborhood factors-enabling interventions that improve care quality while eliminating unnecessary expenses. Hadoop's use of standard hardware made Explorys's storage costs roughly ten times cheaper than traditional relational data warehouses, allowing them to redirect capital toward innovation rather than licensing fees.
NASA demonstrates how even government agencies can harness Big Data through open innovation and crowdsourcing. The NASA Tournament Lab connects real-world challenges with innovative problem solvers through TopCoder's community of software developers, data scientists, and statisticians. When NASA needed applications to explore its 100-terabyte Planetary Data System, they ran a contest seeking ideas rather than assigning employees to the task. The contest attracted 212 registrants and 36 submissions in just two weeks, with the winner proposing a "White Spots Detection" system to identify under-researched areas in planetary systems.
NASA's experience shows that effective Big Data utilization doesn't require enormous rewards-in fact, offering too much money might signal excessive complexity. The most challenging aspect isn't determining which projects to crowdsource but articulating problems in ways that allow diverse solvers to apply their knowledge.
These case studies demonstrate that size doesn't determine Big Data success. Progressive organizations of all types are reaping big rewards by recognizing that Big Data is simply too big to ignore.
Глава 7
Starting Your Big Data Journey: Practical Guidance for Implementation
Before diving into Big Data, organizations must pause and consider several key factors. First, those that don't recognize information's inherent value can't benefit from Big Data. Companies that view data as a problem to minimize rather than an asset to leverage will struggle to extract meaningful insights.
Second, Big Data tools don't cleanse bad data. Organizations with terrible underlying data quality will simply generate inaccurate insights faster. The adage "garbage in, garbage out" applies to all systems, including the most sophisticated Big Data tools.
Third, organizations must consider what they'll do with Big Data short and long-term, what questions they'll ask, whether they're prepared for unexpected answers, and their larger business objectives. If an organization is resistant to data-driven decisions, they should hold off on Big Data initiatives-doing it wrong is worse than not doing it at all.
Finally, while Hadoop and NoSQL databases are freely available, organizations shouldn't confuse "free speech" with "free beer." CIOs who believe they can leverage Big Data without proper budgets, headcount, consulting, or training are mistaken.
When beginning the Big Data journey, start relatively small and organically. Unlike traditional CRM or ERP implementations, companies can begin benefiting from Big Data relatively quickly based on their resources and technical sophistication. While the business case is strong, starting conservatively with reasonable investments is often wiser than boiling the ocean with massive initial expenditures.
Focus first on little victories. Organizations that struggle with Small Data won't accurately predict customer behavior in six months. Instead, target short-term objectives like gathering unstructured customer data or improving website design based on user behavior before tackling more ambitious goals. For maximum impact, start with relatively inefficient business functions where quick wins can build momentum.
Organizations leveraging Big Data need true data scientists, not just repurposed financial analysts. Companies should either invest in training existing employees or aggressively recruit scarce data science talent. Since data scientists won't remain available for long, organizations should act quickly or consider alternatives like consulting firms or crowdsourcing platforms when internal expertise isn't available.
Unlike mission-critical applications like email or ERP, Big Data currently functions as a luxury that organizations can experiment with. Employees should be encouraged to explore new data sources, push boundaries, and embrace the creative, exploratory nature of Big Data work. Like real scientists, data scientists don't follow rigid routines-they swim in data, making unexpected discoveries that yield valuable insights.
Large organizations won't embrace Big Data overnight. As benefits become apparent, internal resistance will decrease. Publicize successes through company wikis, intranets, and external channels like industry journals and conferences. Success stories may even attract data science talent to your organization.
Big Data often reveals counterintuitive insights that challenge conventional thinking. Walmart discovered through data mining that pre-hurricane shoppers bought not just flashlights but also strawberry Pop-Tarts (seven times normal sales) and beer (the top-selling item). To maximize Big Data's value, organizations must remain open to surprising findings.
Unlike traditional relational databases with rigid schemas, Hadoop stores data in its raw form without requiring transformation before loading. This represents a fundamentally more flexible approach to data modeling where meeting business needs trumps following predefined schemas.
Organizations should avoid five common Big Data pitfalls: viewing it as all-or-nothing, treating it as a one-off initiative, assigning it as a side project to already-busy employees, expecting a simple implementation checklist, and allowing IT to "own" it exclusively. Big Data requires embedding data-oriented decision-making into organizational DNA while fostering collaboration between IT and business units.
Глава 8
The Dark Side of Big Data: Ethical Considerations and Challenges
Big Data's vast potential comes with significant ethical questions that, while not new, appear on an unprecedented scale. As Melvin Kranzberg noted, "Technology is neither good nor bad; nor is it neutral"-transformative technologies force us to reconsider established legal, societal, and ethical principles.
Privacy represents the elephant in the room for Big Data. Companies effectively leveraging Big Data often face media scrutiny and public criticism. Google exemplifies this pattern with multiple privacy violations-from Street View software secretly collecting data on open Wi-Fi networks to bypassing Safari privacy settings, resulting in a record $22.5 million FTC fine. Such repeated "privacy-challenged" actions risk government action, lawsuits, brand damage, and consumer exodus.
Beyond privacy concerns, Big Data poses major security challenges. The massive information repositories at companies like Google, Apple (with 400 million customer credit cards), and Amazon represent prime targets for hackers. The bigger the data, the bigger the target. Even smaller players face significant risks-Zappos had to notify 24 million customers after a breach, and LinkedIn suffered the theft of 8 million usernames and passwords, triggering a $5 million lawsuit.
Just as social media experienced a boom followed by user fatigue, Big Data may follow a similar trajectory. Consumers may eventually tire of actively generating data, particularly on social networks. However, unlike social media, Big Data's value won't diminish-early adopters will gain significant competitive advantages by developing customer insights before their competitors catch up.
Big Data threatens to disrupt knowledge workers much as automation eliminated typists, bank tellers, and travel agents. Organizations should expect resistance, particularly from employees who believe their judgment and experience trump data-driven insights. Like doctors who resented patients researching symptoms online, many professionals will resist Big Data because it challenges their authority and forces them to develop new skills.
Mauboussin's research reveals a counterintuitive tendency: as problems grow more complex, many professionals actually rely less on quantitative analysis and more on judgment. This "big paradox" means that when facing increasingly complex data, people often default to simpler, intuitive decision-making-precisely when data-driven approaches might be most valuable.
Supreme Court Justice Potter Stewart's observation that "Ethics is knowing the difference between what you have a right to do and what is right to do" perfectly captures the Big Data dilemma. While Big Data offers tremendous potential to improve business decision-making, it simultaneously amplifies existing ethical issues while introducing new concerns.
Глава 9
The Future of Big Data: From Active to Passive Data Generation
As traditional retailers face unprecedented challenges from online competitors, companies like Target have turned to sophisticated data mining to gain competitive advantage. Target's "pregnancy-prediction model" developed in 2002 proved so accurate that it could identify a pregnant teenage girl before her own father knew of the pregnancy. By analyzing products that potentially pregnant women bought or avoided, Target could send customized coupons to these high-value customers.
Despite ethical controversies surrounding such practices, organizations increasingly recognize Big Data's transformative power. Computer scientist Joe Hellerstein describes this era as "the industrial revolution of data," with the 2012 presidential election potentially marking Big Data's crossing of the chasm into mainstream acceptance.
Several groups are working to make data not just bigger but more democratic, portable, and contextual. The Vibrant Data Project advocates for transparency and asks how data can work for people rather than against them. Google's Data Liberation Front enables users to move their data in and out of Google products, addressing the critical question of data ownership. Meanwhile, the Open Data Foundation is developing metadata standards to establish common understanding across organizations, countries, and languages.
As more people connect to the internet, particularly in developing nations through increasingly affordable smartphones and tablets, humanity will produce exponentially more data. But the real growth will come from devices themselves through the Internet of Things-a fundamental shift where machines gather information without human input.
As Kevin Ashton, who coined the term in 1999, explained: computers need their own means of observing the world "without the limitations of human-entered data." This shift from active to passive data generation will dramatically accelerate data growth. Smart devices like Nabisco's RFID-enabled product tracking can analyze consumer behavior in unprecedented detail-how long shoppers hold products, which packaging elements they focus on, and their nonverbal reactions.
The future promises smart refrigerators that go beyond reading expiration dates to helping people diet and eat better. LG's Smart ThinQ refrigerator includes health manager features that track diets, send recipes to smart ovens, and monitor grocery inventory. Since our data differs, my refrigerator will recommend different dishes than yours.
Beyond the benefits that forward-thinking organizations are realizing from Big Data, there's a more crucial reason for its adoption: survival. As municipalities face dire financial circumstances, Big Data offers solutions for operating more efficiently. Mayor Menino of Boston launched Street Bump not as a publicity stunt but as part of his "New Urban Mechanics" approach to civic innovation. His philosophy-"We are all urban mechanics"-recognizes that government must innovate to survive and provide essential services.
We're living through a permanent data deluge driven by mobility, social web, and IT consumerization. Rather than fight this inevitability, organizations should embrace it to do more with less. While Big Data isn't an elixir for all problems, it's certainly part of the solution. Data has inherent limitations-as Nate Silver notes, statisticians never "bat 1.000" and there will always be "unknown unknowns." Yet this shouldn't stop organizations from beginning their Big Data journey. Those who refuse to integrate data into decision-making will fall behind. As Silver says, "Before we demand more of our data, we need to demand more of ourselves." It all starts with people, not just data and technology.