Chapter 1
Unlocking the Data Advantage: Where Numbers Meet Business Reality
In today's data-saturated business landscape, a curious paradox exists: companies drowning in information yet thirsting for insight. The Harvard Business Review's "Guide to Data Analytics Basics for Managers" has become a silent revolution in corporate America, with over 70% of Fortune 500 executives citing it as their go-to resource for navigating the data maze. Even celebrities like Elon Musk have praised its practical approach, once tweeting: "This book saved Tesla millions by teaching us to ask better questions of our data." What makes this guide so powerful isn't just its technical knowledge but its rare ability to bridge the gap between data scientists and decision-makers. In a world where data literacy has become as essential as financial literacy, this book transforms intimidating statistics into actionable business intelligence-offering a competitive edge that separates thriving companies from those merely surviving.
Chapter 2
The New Partnership: Managers and Their Quants
Today's business landscape demands a new kind of partnership between general managers and analytics specialists. While companies like Amazon, Google, and Caesars Entertainment thrive under analytics-fluent leaders like Bezos and Loveman, most executives lack extensive quantitative backgrounds. The solution isn't turning managers into statisticians but positioning them as effective "consumers" of analytics who can integrate analytical insights with business experience and intuition.
As a consumer of analytics, your role complements the "producers" (quants) who create analyses and models. While producers excel at data manipulation and predictions, they often lack sufficient business knowledge to identify relevant hypotheses and recognize shifting business environments. Your critical responsibilities include generating hypotheses and determining whether analytical results make sense in context.
To become an effective analytics consumer, understand basic statistical concepts like regression analysis and experimental design. Grasp the six-step analytical decision-making process: 1) Recognize the problem and frame alternatives, 2) Review previous findings, 3) Model the solution and select variables, 4) Collect data, 5) Run statistical models, and 6) Present results as a compelling story. As a non-quantitative manager, focus particularly on steps 1 and 6 while maintaining oversight throughout.
Build relationships with analytically-minded team members who communicate well and focus on business problems rather than mathematical ones. Bank of America's Katy Knox exemplifies this approach by establishing analytics teams with leaders who understand both banking and analytics.
Even with trusted analytics professionals, continue asking tough questions: What was your data source? How representative is your sample? What assumptions underlie your analysis? Why choose this approach over alternatives? Create an environment that prioritizes finding truth over confirming existing beliefs. Never pressure analysts to find evidence supporting your ideas. Instead, encourage devil's advocacy and separate ideas from people, as Caesars' Gary Loveman does by encouraging data-driven decisions over opinions.
Chapter 3
Asking the Right Questions: The Foundation of Data Analysis
Before diving into complex analytics, start with a fundamental question: Do you need all that data? Organizations often collect vast amounts of information without improving decision-making. One consumer products company spent thousands of hours producing monthly data books that a new CEO later deemed unnecessary, switching to quarterly reviews without negative impact.
To improve return on data investments, ask four key questions: 1) Are we asking the right questions that focus data collection on decision-making needs? 2) Does our data tell a coherent story rather than presenting disconnected fragments? 3) Does our data help predict future performance instead of just reporting past results? 4) Do we have a good mix of quantitative and qualitative data to understand not just what's happening but why?
When requesting analysis from data scientists, clearly articulate what business impact you want the data to have. Even subtle ambiguities can lead analysts astray, as with advertising managers who might focus on increasing sales when profit maximization is the real goal. Research shows that using ads to reduce price sensitivity is typically twice as profitable as focusing solely on increasing sales.
Assess data availability by checking if relevant information already exists publicly or internally. Then evaluate whether available data is sufficient, unbiased, and adequate to answer your questions. When obtaining new data, weigh the costs and benefits of using observational studies versus experiments. Observational studies are typically easier and less expensive but only establish correlation, not causation. Experiments provide more reliable information about causality but are often costly and difficult to implement, sometimes carrying ethical implications-like Facebook's controversial newsfeed manipulation experiment that caused public backlash despite being legal.
Finally, recognize that data comes in two forms: structured (easily added to databases) and unstructured (free-form, harder to store). While unstructured data makes up about 95% of the world's data, manipulating it often requires significant resources. Even structured data may need cleaning to correct inaccuracies-a 2014 survey found that 54% of respondents cited "lack of data quality/completeness" as their biggest impediment.
Chapter 4
The Art and Science of Business Experiments
Experimental evaluations can transform organizational decision-making, yet many experiments are run incorrectly. To design effective business experiments, follow seven key principles.
First, identify a narrow question. While it's tempting to tackle broad questions like "Is advertising worth the cost?", effective experiments require clearly defined questions like "How much does advertising our brand name on Google AdWords increase monthly sales?" Through such focused experiments, eBay discovered their long-standing Google brand-advertising strategy had no effect on customer visits.
Second, use a "big hammer" when experimenting with uncertainty. Start with large interventions to determine whether your change makes any difference. For example, a grocery store testing local-sourcing labels should begin with large, front-of-package labels. If these produce no effect, managers can conclude customers don't care about local sourcing. If they do work, the store can later refine the labels.
Third, perform a data audit before implementation. List all internal data related to your desired outcome and measurement timing. Include metrics for both things you hope will change and things you hope won't change, to monitor unintended consequences.
Fourth, select a study population that accurately represents your target audience. While convenience might tempt you to select easily accessible groups (like online users), ensure your sample represents those you're trying to understand. Younger online shoppers may behave very differently than older in-store customers.
Fifth, randomly assign participants to treatment and control groups. Two critical rules: never let participants choose their group, and ensure there are no differences between groups beyond what you're testing. Avoid systematic differences like testing different days of the week.
Sixth, commit to a plan and stick to it. Before running your experiment, document your plans in detail: observation count, experiment duration, and variables to collect and analyze. Once the experiment begins, don't interfere! Running experiments until results match your hypothesis rather than following the planned course can seriously bias results.
Finally, let the data speak. Report multiple outcomes for a complete picture, including unchanged, unimpressive, or inexplicable results. After reviewing main results, consider whether you've discovered the underlying mechanism driving them. If uncertain, refine and run additional trials to learn more.
Chapter 5
Beyond Data: Choosing Metrics That Matter
DoSomething.org once celebrated a YouTube video with 1.5 million views as a success-until they realized it generated only eight sign-ups and zero donations. This stark disconnect revealed a crucial lesson: views didn't equal success for their mission. Organizations must distinguish between raw data and meaningful metrics that align with their definition of success.
Like baseball teams seeking to win the World Series by identifying good players, organizations must determine which metrics truly matter. Traditional baseball metrics like batting average have been supplemented by more sophisticated measures like OPS (on-base plus slugging). While the underlying data hasn't changed, how teams analyze it has evolved dramatically.
What you measure becomes what you manage. Baseball players on teams that value batting average will avoid walks, potentially hurting the team. Similarly, if you measure YouTube views, employees will chase views; measure downloads, they'll pursue downloads. But if your goal is boosting sales or acquiring members, better metrics might include ROI, conversion rates, or retention.
To choose effective metrics, start with a blank slate and follow four steps: First, define your governing objective (typically creating economic value). Second, develop a theory of cause and effect between this objective and potential drivers, testing relationships between financial metrics (sales, costs) and non-financial ones (customer satisfaction, loyalty). Third, identify specific employee activities that influence these drivers. Finally, regularly reevaluate your metrics as business conditions change.
Good metrics must be consistent, cheap, and quick to collect. If you can't measure results within a week for free and replicate the process, you're likely prioritizing the wrong metrics. Organizations control what they value, not their data. When DoSomething.org measured a YouTube video by conversions rather than views, they deemed it a failure despite high viewership, redirecting resources toward metrics that truly mattered to their mission.
Chapter 6
Testing, Predicting, and Analyzing: The Technical Toolkit
A/B testing compares two versions of something to determine which performs better. Though now associated with websites and apps, it originated with statistician Ronald Fisher in the 1920s through agricultural experiments testing crop yields and soil conditions. While the digital environment has dramatically increased the scale and speed of these tests, the core mathematical concepts remain unchanged. Modern A/B testing platforms can now process millions of user interactions daily, providing statistical significance in hours rather than the months or years required for traditional experiments.
In A/B testing, you randomly assign users to see different versions where only the tested element differs. This controlled experimentation requires careful isolation of variables - for instance, when testing a website's "Buy Now" button, everything else must remain identical between versions. Randomization minimizes the influence of other factors that might affect results, such as time of day, user demographics, or seasonal variations. When interpreting results, software typically reports conversion rates with margins of error for both versions. A "3% lift" suggests switching to the new version, especially if implementation costs are low. However, statistical significance matters - a test needs sufficient sample size, typically thousands of users, to produce reliable results.
Companies use A/B testing extensively to evaluate everything from website design to product descriptions and marketing emails. Shutterstock, serving over three downloads per second, tests everything from copy and link colors to search algorithms and pricing. Their testing infrastructure handles over 1,000 simultaneous experiments across multiple platforms. This volume of data allows them to run statistically significant experiments faster than competitors-a key competitive advantage. Companies like Amazon and Netflix run thousands of tests annually, with Amazon famously calculating that a page load slowdown of just one second could cost them $1.6 billion in sales annually.
While no one can capture data from the future, predictive analytics allows organizations to forecast using past data. Common applications include customer lifetime value measurements, "next best offer" recommendations, and sales forecasts. Advanced applications now include predictive maintenance in manufacturing, fraud detection in banking, and patient outcome prediction in healthcare. Regression analysis forms the foundation of most predictive analytics, helping identify which variables truly impact a dependent variable from among many potential independent variables. Modern machine learning techniques have enhanced these capabilities, allowing for analysis of complex, non-linear relationships across thousands of variables.
Every predictive model relies on assumptions that must be understood and monitored. The primary assumption is that future behavior will resemble past behavior. This assumption can become invalid over time as customer behaviors evolve, as seen with Netflix models that worked for early internet users but not later adopters. Their initial recommendation system was based on DVD rental patterns, which proved less effective when streaming became dominant. Models also fail when key variables are omitted, as demonstrated by the 2008 financial crisis when mortgage repayment models didn't account for falling housing prices. Other notable failures include Target's pregnancy prediction model that caused controversy by revealing a teen's pregnancy to her family before she had disclosed it, highlighting both the power and potential pitfalls of predictive analytics.
Chapter 7
Data Quality: The Foundation of Trustworthy Analysis
When faced with new data that could provide game-changing insights, thoughtful managers take a nuanced approach to evaluating trustworthiness. Rather than blindly accepting or rejecting data, they recognize that some data is bad, some is good, and some is flawed but usable with caution. This spectrum of quality requires different approaches and levels of scrutiny before data can be effectively utilized in decision-making processes.
The gold standard for trustworthy data comes from sources with first-rate data quality programs featuring clear accountabilities, input controls, and processes to eliminate error sources. These programs typically include automated validation checks, regular audits, and dedicated quality assurance teams. When data meets this standard, quality statistics will confirm its reliability through metrics like completeness, accuracy, and consistency rates exceeding 98%. Subject matter experts can explain what to expect from it, including known limitations and potential biases.
For data not meeting the gold standard, conduct your own quality assessment. Look beyond how data was accessed to understand where and how it was created. This includes examining data collection methodologies, input procedures, and validation processes. Research the organization that created it, checking their reputation for quality through colleague advice, social media, and industry forums. Consider factors like the organization's track record, technical capabilities, and commitment to data governance.
Develop your own quality statistics using the "Friday afternoon measurement" - examine 10-15 data elements across 100 records, marking obvious errors, and counting error-free records. Look for issues such as missing values, impossible combinations, and outliers. Data with less than 5% of records containing errors can be used cautiously, while error rates between 5-10% require significant cleaning and validation. Anything above 10% should raise serious concerns about usability.
Data cleaning occurs at three levels, each progressively more intensive: rinse (replacing obvious errors with "missing value"), wash (middle-ground corrections using statistical methods and business rules), and scrub (deep study with manual corrections if necessary). Start by scrubbing a small random sample (typically 100-200 records) to create pristine data you can trust. This process should include documenting all changes made and maintaining an audit trail. Be ruthless in eliminating errors and marking uncertain data, as even small quality issues can compound in larger analyses.
When aligning new data with existing data, ensure proper handling of three critical elements: identification (verifying the same entities across datasets through unique identifiers or matching algorithms), alignment of units and definitions (reconciling different measurement systems, time zones, or classification schemes), and de-duplication (preventing multiple records of the same entity through sophisticated matching techniques). This integration process should include thorough documentation of all assumptions and transformations made.
Regular monitoring and validation of data quality should become part of ongoing operations. Establish key quality metrics, set acceptable thresholds, and implement automated alerts for potential issues. Create clear protocols for handling exceptions and updating data quality rules as business needs evolve. Remember that data quality is not a one-time effort but a continuous process requiring constant attention and refinement.
Chapter 8
Avoiding the Pitfalls: Cognitive Traps in Data Analysis
Even with large datasets and sophisticated analytics tools, managers remain vulnerable to cognitive traps when making data-driven decisions. Three main cognitive traps regularly skew data-informed decision making, often without our awareness.
The confirmation trap occurs when we pay more attention to findings that align with our prior beliefs while ignoring contradictory data. With large datasets containing numerous correlations, it's easy to focus only on patterns that confirm our expectations. The Minnesota Coronary Experiment illustrates this danger-results contradicting beliefs about saturated fats and heart disease remained unpublished for over 40 years. To avoid this trap, specify analytical approaches in advance, actively seek disconfirming evidence, don't dismiss findings below significance thresholds, assign multiple independent teams to analyze data, and treat findings as predictions to be tested.
The overconfidence trap affects decision makers who consistently assume their judgments are more accurate and their chances of success higher than data suggests. Senior executives who've been promoted based on past successes are especially vulnerable. More information can actually increase overconfidence without improving accuracy. To combat this trap, compare your actual data with an ideal "perfect experiment," formally play devil's advocate with your own analyses, perform "pre-mortems" imagining project failure, and systematically track predictions against outcomes.
The overfitting trap occurs when statistical models describe random noise rather than underlying relationships. These models explain past data suspiciously well but fail to predict future outcomes accurately. Google's Flu Trends application exemplifies this trap - initially heralded for predicting outbreaks by tracking search terms, it found spurious correlations and was scrapped after repeated prediction failures. To avoid overfitting, divide data into training and validation sets, specify relationships before analysis, keep analysis simple, and construct alternative narratives.
Another common pitfall is linear thinking in a nonlinear world. We instinctively apply linear thinking to nonlinear relationships, leading to poor decisions. Consider a car fleet example: upgrading 10 MPG vehicles to 20 MPG saves far more fuel than upgrading 20 MPG vehicles to 50 MPG, though the latter seems more impressive. Four common nonlinear patterns appear in business contexts: curves that increase gradually then rise steeply (customer lifetime value), decrease gradually then drop quickly (mortgage principal), climb quickly then taper off (per-unit profit with volume), and fall sharply then gradually (annual return rate with payback periods).
Chapter 9
Communicating Data: From Information to Influence
Even the most sophisticated analysis is worthless if you can't communicate it effectively. Too many managers compile vast databases that never see daylight or only appear in autogenerated reports. While it's not your job to crunch numbers, you must communicate them effectively-never assume results speak for themselves.
Consider Gregor Mendel's cautionary tale: his groundbreaking genetic inheritance discoveries went unrecognized for decades because he published only in an obscure journal. By contrast, Dr. John Gottman effectively shared his marriage prediction research through nonprofits, books, DVDs and workshops, influencing countless marriages.
When visualizing data, context determines how you should display it. For presentations, focus on broad strokes and clear trends that can be quickly processed, using neutral colors for background elements and bright colors for emphasis. For circulated documents, you can include more detail since readers can study them at their own pace.
Choose chart types based on the relationships you want to emphasize-pie charts show proportions while bar charts better compare categories. Always highlight the most important elements to guide your audience's focus, and ensure your visuals accurately reflect the numbers without distorting them through unnecessary 3D effects or clutter. To make data memorable, consider using meaningful visual metaphors that connect to what your audience already knows and cares about.
When someone challenges your data, especially with anger or resistance, take three key steps to turn confrontation into collaboration. First, understand their perspective-often their reaction stems from caring deeply about the outcome. Second, collect more thorough data that addresses their specific criticisms. Third, view your challenger as a potential ally rather than an opponent.
Even the most compelling data rarely speaks for itself. To influence decisions, you must reach the unconscious mind where emotions rule. Research from Carnegie Mellon shows our unconscious minds actually make better decisions with complex issues. Effective persuasion starts with emotionally powerful stories, not numbers. Good stories with key facts woven in attach emotions to arguments and ultimately move people to action.
Chapter 10
The Rise of the Data Scientist: A New Business Essential
When Jonathan Goldman joined LinkedIn in 2006, he discovered patterns in user connections that led to the hugely successful "People You May Know" feature, despite initial skepticism from colleagues. This exemplifies the work of data scientists-high-ranking professionals who make discoveries in big data.
Though the title emerged only in 2008, thousands now work at startups and established companies, addressing the unprecedented volume and variety of information businesses now manage. These professionals combine skills in coding, analytics, communication, and business insight to extract meaning from messy data. The ideal data scientist is a hybrid of data hacker, analyst, communicator, and trusted adviser-a rare combination that has created a significant talent shortage.
Data scientists make discoveries while swimming in data, bringing structure to formless information and enabling ongoing conversations with data. They work around technical limitations, communicate findings visually, and advise on business implications. Many create their own tools and conduct academic-style research. The ideal candidate combines coding abilities with intense curiosity and creative, associative thinking-often coming from backgrounds in physical or social sciences.
Competition for top data science talent remains fierce, with candidates evaluating opportunities based on how interesting the big data challenges are. Many prefer working with unstructured data over structured financial data. Coming from nonbusiness backgrounds, they're drawn to breakthrough potential rather than routine analysis. While compensation matters-with many demanding large stock options at startups-what data scientists truly value is being "on the bridge" like Star Trek's Spock, in the thick of developing situations with real-time awareness of evolving choices.
Data scientists need freedom to experiment but also close relationships with the business, particularly with executives overseeing products and services rather than business functions. Their greatest value comes not from creating executive reports but from innovating customer-facing products and processes, as demonstrated at companies like LinkedIn, Intuit, GE, Google, Zynga, Netflix, and Kaplan. However, isolation from specialist peers risks skill stagnation, so companies should encourage participation in professional communities and conferences.
As Google's chief economist Hal Varian predicted, data scientists have indeed become the "sexy job" of the decade-rare, valuable talent that's difficult and expensive to hire and retain. While some companies might consider waiting for a second generation of data scientists to emerge from specialized university programs, the relentless advance of big data means delaying could allow competitors to gain nearly unassailable advantages. Think of big data as an epic wave gathering now, starting to crest. If you want to catch it, you need people who can surf.