Inside the Science behind Smarter Marketing Measurement with Professor Ron Berman

Guest: Professor Ron Berman (Associate Professor of Marketing, The Wharton School; Scientific Adviser to Lifesight) | Host: Rajeev Nair | Podcast: Humans of Measurement.

Organizations today have access to more dashboards, attribution models, and data points than ever before. Yet, generating more metrics does not automatically translate to better decision-making.

In this episode of Humans of Measurement, Rajiv Nayer speaks with Professor Ron Berman, Associate Professor of Marketing at the Wharton School and Scientific Adviser to Lifesight. Drawing from his background across software engineering, venture capital, and academic marketing science, Professor Berman breaks down why precision without purpose is useless, the hidden prevalence of false discovery rates in A/B testing, why marketing lags product teams in experimentation, and how to properly right-size experiments depending on whether you are making a quick choice or calibrating a long-term model.

Key Takeaways

  • The Precision Trap: Data science teams often obsess over making metrics more precise, but roughly 80% cannot explain what specific business decision a more precise metric actually enables.

  • The 20% False Positive Reality: In standard A/B testing (at the p < 0.05 significance level), when companies run dozens of minor tests with low underlying true impact, roughly 20% of “winning” results are actually false discoveries.

  • When “Hacking” Experiments Works: You do not always need a massive sample size. If your goal is simply choosing between Option A and Option B (“Test and Roll”), smaller sample sizes suffice. However, if you are calibrating a Marketing Mix Model (MMM) or establishing a baseline, rigorous, fully powered experiments are required.

  • Why Marketing Lags Product in Experimentation: Unlike product feature tests—which cost zero ad spend and carry permanent effects—marketing experiments incur real media costs, face changing seasonality, and carry public reputational risks.

  • AI and “Decision-Driven Analytics”: The future of measurement lies in using AI to test dynamic decision frameworks, reconciling conflicting models (e.g., MMM showing 20% ROI vs. Attribution showing 30% ROI), and working backward from the decision to the data.

Core Discussion Topics

1. The Precision Trap: Why More Metrics Don’t Mean Better Decisions

Data science and analytics teams naturally strive for precision—they want unbiased, highly accurate metrics. However, Professor Berman points out a disconnect between number generation and management decision-making.

When teams request “more precise measurement,” the first question to ask is: Why? What decision will this enable us to take that we couldn’t take before? In most cases, teams cannot give a clear answer. Generating a more precise number does not automatically improve outcomes unless it crosses a threshold that changes a specific business action. High-performing marketing organizations distinguish themselves by identifying exactly when precision matters for a decision and when a rougher estimate is sufficient.

2. False Discovery Rates (FDR) in A/B Testing

In standard hypothesis testing, statistics guarantees that false positives occur only 5% of the time (p < 0.05). However, that rule answers: “Among all experiments with zero true effect, how often will I get a false positive?”

What managers actually care about is: “Given that I got a statistically significant result, what is the chance it is a false positive?”

When companies run thousands of minor A/B tests (e.g., changing button colors or headline tweaks) where the true underlying effect is zero or negligible, the probability of a false positive among the “successful” tests skyrockets. In a study analyzing over 2,700 experiments on Optimizely, Professor Berman found that roughly 20% of statistically significant results were actually false discoveries. Without controlling for False Discovery Rates (FDR), companies frequently deploy changes that have no real impact.

3. “Test and Roll” vs. Fully Powered Experiments

Vendors and growth teams often push for fast, two-week “hacked” experiments. While some measurement providers reject this practice, Professor Berman offers a nuanced perspective:

  • When to use short/small experiments (“Test and Roll”): If an experiment’s sole purpose is to pick Option A over Option B (e.g., choosing between two creative variants for a short campaign) and the outcome won’t be used for long-term modeling, waiting for massive sample sizes carries a high opportunity cost. If the true performance difference between A and B is tiny, picking the wrong one doesn’t harm the business significantly.

  • When to use rigorous, fully powered experiments: If you are running an experiment to establish long-term baseline learning, measure exact lift (e.g., A is better than B by exactly 11% ± 2%), or calibrate a cross-channel Marketing Mix Model (MMM), “hacking” the test fails. These use cases require strict, fully powered statistical controls.

4. Why Marketing Lags Product Teams in Experimentation

Product teams in tech companies adopt experimentation rapidly through Minimum Viable Products (MVPs) and feature flags, whereas marketing teams adopt experimentation much more slowly. Professor Berman highlights four structural reasons for this gap:

  1. Cost and Lift Dynamics: A website feature test costs zero additional media spend and can yield a massive lift. An ad campaign test requires real media budget, while individual marketing interventions often produce subtle, hard-to-detect lifts that demand giant sample sizes.

  2. The Observational Data Trap: Marketing teams are swimming in historical tracking and attribution data, creating a strong temptation to analyze existing data rather than pausing spend to run clean experiments.

  3. Lack of Decision Continuity: In product development, once a feature wins an A/B test, it is permanently deployed. In marketing, channel dynamics shift constantly due to seasonality and competitor behavior; showing that Google beat Meta this week does not guarantee the same budget allocation applies next month.

  4. Reputational Risk: A broken website feature can be reverted in hours, but bold, public marketing experiments carry brand reputation risks that make managers risk-averse.

5. Reconciling Conflicting Models and the Future of AI

When an MMM shows a channel yielding a 20% ROI while a Multi-Touch Attribution (MTA) model shows a 30% ROI, managers often panic and assume one model is broken. Professor Berman notes that neither model is inherently “wrong”—they are measuring different aspects of the customer journey using different assumptions. The frontier of marketing science lies in developing frameworks that combine conflicting models into a unified decision engine.

Looking toward AI, Professor Berman suggests that large enterprise organizations can begin running A/B tests between AI agents and human managers—allocating a portion of budget to an automated agent and evaluating revenue and acquisition cost performance against human-managed campaigns.

Quote of the Episode

“When teams tell me, ‘I want more precise measurement,’ the first thing I ask them is: ‘Why? What will a more precise measurement enable you to do, and what decisions can you take?’ In my experience, 80% of the time, I don’t get an answer… Identifying when precision matters for a decision is what defines a strong marketing organization.”

— Professor Ron Berman

Actionable Steps for Measurement Teams

  1. Apply the Decision-First Test: Before requesting or building new measurement models, document the exact business decision the model will inform and the threshold required to change your budget allocation.

  2. Account for False Discoveries: When running large volumes of low-impact A/B tests, implement False Discovery Rate (FDR) controls to avoid rolling out changes that offer zero true incremental lift.

  3. Match the Test to the Intent: Use quick, smaller sample size tests when choosing between immediate creative variants (“Test and Roll”), but enforce strict, fully powered experimental holds when calibrating cross-channel MMMs or establishing true baselines.

  4. Leverage AI for Fast Modeling: Use modern AI assistants to brainstorm mathematical formulations and run rapid simulations to validate measurement hypotheses before devoting weeks of engineering resources.

About the Guest

Professor Ron Berman is an Associate Professor of Marketing at The Wharton School of the University of Pennsylvania and serves as a Scientific Adviser to Lifesight. His academic research spans digital marketing, experimentation, attribution, and AI. Prior to academia, he worked as a software developer and worked at a venture capital firm analyzing go-to-market strategies.

Have questions or want a personalized demo?

You May Also Like

  • Beyond Attribution cover

    Beyond Attribution: Causality & Incrementality with Professor Kenneth Wilbur

    Beyond Attribution: Causality & Incrementality Guest: Professor Kenneth Wilbur, Professor of Marketing and Analytics, UC...

  • Everyone Blames Marketing with Lara Schoisman cover

    Everyone Blames Marketing with Lara Schoisman

    Everyone Blames Marketing: Execution, AI, and Owning Your Niche Guest: Lara Schoisman, Founder of...

  • The Missing Discipline in Performance Marketing with Mark Friedman cover

    Making Marketing Make Sense to the CFO with Scott Davidson

    Making Marketing Make Sense to the CFO | Scott Davidson on Attribution, AI & Marketing...