Million-dollar decisions are being made by businesses across functions relying on forecasts without ever seeing how the model behind them was chosen, or how often it has been right.
We looked at that decision. The common approach is to fit a set of models, score them, and keep the winner for each series. It’s sensible, it’s what most teams do, and as far as we could find it hadn’t been tested out-of-sample at scale on real advertiser data. So we ran the comparison and published what came back.
Everything below is what we saw on this data.
What we found
Last quarter’s best model probably won’t be this quarter’s: No model was best more than about a third of the time, and the rankings barely held from one origin to the next. With no stable winner to find, scoring models per series bought almost nothing over picking one model and leaving it alone.
Combining models did as well as picking the right one: The ensemble matched what you’d get from choosing the single strongest model and applying it everywhere, except you never have to work out which model that is. Against picking a different model for each series it was clearly ahead, in every industry we looked at.
Fewer misses against the plan: The measure we found most useful is regret: how much accuracy a strategy gives up against the best model that was actually on the table. That gap is what tends to surface later as a miss against plan. The ensemble carried the lowest regret of anything we tested, on a typical forecast and on the worst one.
The reason appears to be diversity. The models don’t miss in the same direction at the same time, and averaging turns that disagreement into accuracy.
Useful from the first forecast: One selection rule did keep up: pooling validation error across every series and applying one model everywhere. It matched the ensemble on typical accuracy, but only after a fair amount of backtest history had built up. Before that it often landed on the weakest model in the panel. The ensemble has no ranking to estimate, so it works from the first origin, which matters most for advertisers with the least history behind them.
Fewer bad quarters: A well-chosen single model was outright best slightly more often than the ensemble. It also landed near the bottom of the pack far more often, and that’s the part a portfolio feels. One bad quarter a year on a portfolio-wide forecast means a re-forecast and a credibility problem.
The trade is cost against robustness: Running one model is cheaper. Running several and combining them costs more, and what the extra cost buys is robustness. Models drift, and when a single model starts to degrade it passes that straight through to every forecast you publish. An ensemble absorbs it, because the other members are still pulling the average back toward where it should be.
What we’d take from this
Ensembling looks like a reasonable default on the data we tested. Pooled portfolio-level selection becomes a fair alternative once a panel is mature and one model is clearly ahead. Per-series selection is the one we’d be most cautious about.
How we ran it
We ran daily commercial series from dozens of organizations, five years of history, across retail, apparel, food and drink, home and garden, and beauty and fitness. The same media spend covariates for every model, rolling origins, out-of-sample throughout. Every decision was made on data that ended before the forecast began.
The panel skews toward larger, well-instrumented advertisers, since we screened for long histories with near-complete spend coverage. On messier data, knowing which model to pick is harder still.
Scientifically Reviewed
This work is reviewed by our independent Scientific Advisory Council, a panel of experts who assess our research at arm’s length from the product teams.
The data, every forecast we produced and the code that scored them are published alongside the paper, so the result can be checked rather than taken on trust. We’re also open-sourcing Horizon, the engine behind this work, which carries the same model families and combination logic we tested.
If you’d like the full method and results, you can read the paper here.
You may also like
Essential resources for your success








