IN THIS ARTICLE

Million-dollar decisions are being made by businesses across functions relying on forecasts without ever seeing how the model behind them was chosen, or how often it has been right.

We looked at that decision. The common approach is to fit a set of models, score them, and keep the winner for each series. It’s sensible, it’s what most teams do, and as far as we could find it hadn’t been tested out-of-sample at scale on real advertiser data. So we ran the comparison and published what came back.

Everything below is what we saw on this data.

What we found

Last quarter’s best model probably won’t be this quarter’s: No model was best more than about a third of the time, and the rankings barely held from one origin to the next. With no stable winner to find, scoring models per series bought almost nothing over picking one model and leaving it alone.

01 no model stays ahead - Lifesight

Combining models did as well as picking the right one: The ensemble matched what you’d get from choosing the single strongest model and applying it everywhere, except you never have to work out which model that is. Against picking a different model for each series it was clearly ahead, in every industry we looked at.

02 head to head error - Lifesight

Fewer misses against the plan: The measure we found most useful is regret: how much accuracy a strategy gives up against the best model that was actually on the table. That gap is what tends to surface later as a miss against plan. The ensemble carried the lowest regret of anything we tested, on a typical forecast and on the worst one.

03 regret - Lifesight

The reason appears to be diversity. The models don’t miss in the same direction at the same time, and averaging turns that disagreement into accuracy.

Useful from the first forecast: One selection rule did keep up: pooling validation error across every series and applying one model everywhere. It matched the ensemble on typical accuracy, but only after a fair amount of backtest history had built up. Before that it often landed on the weakest model in the panel. The ensemble has no ranking to estimate, so it works from the first origin, which matters most for advertisers with the least history behind them.

04 works from day one - Lifesight

Fewer bad quarters: A well-chosen single model was outright best slightly more often than the ensemble. It also landed near the bottom of the pack far more often, and that’s the part a portfolio feels. One bad quarter a year on a portfolio-wide forecast means a re-forecast and a credibility problem.

05 fewer bad quarters - Lifesight

The trade is cost against robustness: Running one model is cheaper. Running several and combining them costs more, and what the extra cost buys is robustness. Models drift, and when a single model starts to degrade it passes that straight through to every forecast you publish. An ensemble absorbs it, because the other members are still pulling the average back toward where it should be.

06 failure shape 1 - Lifesight

What we’d take from this

Ensembling looks like a reasonable default on the data we tested. Pooled portfolio-level selection becomes a fair alternative once a panel is mature and one model is clearly ahead. Per-series selection is the one we’d be most cautious about.

How we ran it

We ran daily commercial series from dozens of organizations, five years of history, across retail, apparel, food and drink, home and garden, and beauty and fitness. The same media spend covariates for every model, rolling origins, out-of-sample throughout. Every decision was made on data that ended before the forecast began.

The panel skews toward larger, well-instrumented advertisers, since we screened for long histories with near-complete spend coverage. On messier data, knowing which model to pick is harder still.

Scientifically Reviewed 

This work is reviewed by our independent Scientific Advisory Council, a panel of experts who assess our research at arm’s length from the product teams.

The data, every forecast we produced and the code that scored them are published alongside the paper, so the result can be checked rather than taken on trust. We’re also open-sourcing Horizon, the engine behind this work, which carries the same model families and combination logic we tested.

If you’d like the full method and results, you can read the paper here.

Read the full paper →

Rajeev Nair

Rajeev Nair  Linkedin Logo

Rajeev Nair is the Co-Founder of Lifesight, where he plays a key role in building a data-driven marketing measurement platform that helps brands optimize performance and drive growth. With deep expertise in analytics and technology, Rajeev focuses on enabling businesses to make smarter, insight-led decisions.

You may also like

  • ChatGPT ads X Lifesight

    Published on: August 19, 2026

    Introducing ChatGPT Ads Integration in Lifesight

    ChatGPT Ads is the first ad channel to run inside an LLM, and the brands that measure it will be the ones who started measuring on day one. 

  • Mia Blog Thumbnail

    Published on: March 18, 2026

    Introducing MIA, agents that turn measurement into action

    Marketing does not lack insights. It lacks decision velocity to act on them. MIA accelerates it.

  • Going Beyond the Obvious How Advertisers are Using Lifesight Audiences to Improve Engagement by 200x b2a5fbea62 1 - Lifesight

    Published on: October 16, 2023

    Going Beyond the Obvious: How Advertisers are Using Lifesight Audiences to Improve Engagement by 200x

    Discover precision targeting and lookalike segments that drive results. Reinvent your marketing strategies for better ROI. Learn More

Essential resources for your success

  • Measure Drives Growth

    Annual Causal Measurement Report

    Global advertising spend hit $1.14 trillion in 2025. Upto 47% of it is wasted due to poor measurement. Only 52% of CMOs can prove marketing's financial impact.

  • A Guide To Marketing Effectiveness Measurement For Ecommerce Brands

    A Guide To Marketing Effectiveness Measurement For Ecommerce Brands

    Turn fragmented data into clear insights that improve ecommerce marketing performance.

  • Marketing Measurement

    Mastering the Four Pillars of Marketing Measurements

    Learn how each pillar plays a unique role in measuring marketing effectiveness and improving ROI across channels.