We built a marketing mix model (MMM, a model that splits sales between media channels and everything else) for a ski resort's revenue. In the search for the best version we made 886 model runs, counting repeat runs used as checks. The validation score was R-squared, measured on dates held back from the model: 0 means no better than guessing the average every time, 1 means every rise and dip predicted, and below 0 means worse than that guess. The best single score was 0.5243, on a Christmas period the model had not seen: about half the rise and fall in revenue over those dates.

The search's five finalists were also scored a stricter way. Each was trained on earlier dates, asked to forecast the stretch that followed, then moved forward and tested again: 14 test periods in all. Scored across all 14 together, all five fell below 0, from -0.31 to -0.57: their errors added up to 31% to 57% more than simply guessing the average. A strong score on one period can hide a model that fails once every period is counted.

In the team's placebo checks, a fake channel that should earn nothing took a large share of the credit while the score barely moved. That led the team to conclude that the Christmas score came from the model's seasonal baseline, the part that follows the ski season, not from media. These ski-resort results come from one search (a different round of testing from the model in our article on channel credit).

For the best finalist, two of the 14 periods produced 53% of all its error, measured so that big misses weigh the most. Without those two, its score rose only to +0.13, about 13% of the rise and fall, far below the R-squared above 0.8 the contract asked for. So we did not ship a model: as the contract allowed, we delivered a memo instead. It said this data could not reliably show how efficient each channel was, and that a real answer would need experiments, such as regional tests, not more tuning of the model's starting assumptions.

Key Takeaways

Trust an MMM's validation score only if the model was tested on later dates it never saw and beat repeating last year's numbers on those same dates.

  • A ski-resort model's best single score was 0.5243, about half the rise and fall, on a Christmas period it had not seen, but the team concluded the season earned it, not media. Over all 14 test periods, every finalist did worse than guessing the average.
  • On a household-goods brand, a causal-effect method (DoubleML) gave channel effects 1.50 to 2.49 times apart depending only on whether the weeks were split into random groups or unbroken blocks, with no error or warning.
  • In a separate campaign at the same ski resort, repeating last year's numbers beat every model in that comparison in five of six revenue segments and tied in the sixth.
  • Before you move budget, ask for the exact test dates, the last-year comparison on the same dates, and each test period's share of the error.
Ski-resort revenue model: best single period positive, every finalist below zero over all 14 test periods Best single score, Christmas: +0.5243 Best finalist, all 14 periods: -0.31 Worst finalist, all 14 periods: -0.57 worse than guessing better
Ski-resort revenue model, R-squared on dates it had not seen (1 = every rise and dip predicted, 0 = no better than guessing the average, below 0 = worse). The Christmas period alone looked strong; over all 14 test periods, every finalist fell below zero.

Was the model tested on dates it could not have seen?

It should be: trained on earlier dates, then a gap, then tested on the dates that follow.

A fair test works like real life. A model trained on January to October must forecast November without having seen anything from November: not its sales, not the weeks after it, and no rescaling worked out with November's numbers included. If any of that slips in (analysts call it leakage), the score shows how much the test gave away, not how well the model forecasts.

Even a clean score shows only that the forecast holds. Whether each channel's credit is right needs a separate check, such as a placebo channel.

Random cross-validation splits the weeks at random into a set number of groups, hides one group at a time and trains on all the others, including most of the weeks right next to each hidden one. Neighboring weeks look alike, so the model can partly fill in a hidden week from its neighbors instead of forecasting it.

A blocked split hides a run of consecutive weeks instead, which cuts down that effect, but it can still train on weeks after the hidden run. Those later weeks still carry the effect of ads that ran during the test, because an ad keeps working for a while after it runs (modelers call this carry-over, or adstock).

A rolling-origin split goes further: it trains only on earlier weeks and tests the weeks that follow. Validation code we later built for the ski-resort client also leaves a gap between the two, called an embargo. The gap is at least 28 days, and never shorter than the longest carry-over the model allows for any channel.

A fair test: learn from earlier dates, leave a gap, test what follows Train Earlier dates only Rescaling learned here only Embargo (a gap) At least 28 days Covers the longest ad carry-over Test The next dates, in order Season-only forecast scored too
That code also refuses to run if a training period overlaps its test period or the dates are out of order, and will not score a model on fewer than four test periods.

How the weeks are split can change the answer itself, not just the score. For a household-goods brand with 89 weeks of national data, we ran a DoubleML audit (double machine learning, a method that estimates how each channel's spend moves the brand's weekly clicks through to retailers after setting aside what trend and season explain). The method splits the weeks into groups along the way, and we changed only how: at random (shuffled), or in unbroken blocks.

That one switch changed Facebook's estimated effect by a factor of 1.90, TikTok's by 2.49 and Google Ads' by 1.50, and nothing errored or warned. If the split did not matter, each factor would be close to 1. Measuring the channels by impressions instead of spend gave the same pattern, factors of 1.79 to 1.94.

This is real client data, so nobody knows the true effects or which set is nearer the truth. We only know that the shuffled version let the method learn from the weeks directly on either side of each week being estimated, which the blocked version mostly avoided. Because of that swing, the client's report carried no effect sizes from this method, only a plain verdict for each channel.

What else can make the score misleading?

Four things: rescaling on all the data at once, test dates reused for training, picking the best of many runs, and how the test dates are cut.

Rescaling on all the data at once. Many models rescale each input before fitting, for example relative to a channel's largest week. If that is worked out on the full history before the split, a spike in a test week changes how every earlier week is measured. No test week is used for training, yet its information reaches the model. The rule: anything that learns from the data, including rescaling, filling in missing values, choosing which inputs to keep and removing season before media is measured, must learn from the training dates only.

Test dates reused for training. In another project, two test periods did not share a single day, so they looked like two separate checks. But the model tested on the second period had been trained on all of the first period's test dates. That is normal when forecasting forward step by step, but the two results are not independent confirmations and should not be counted as two. Ask for every test period's training and test dates.

Picking the best of many runs. Make hundreds of model runs and report only the best score, and part of that score is luck. It is the same reason the winning ad out of hundreds of variants tends to lose some of its lift when you run it again. The ski resort's 0.5243 was the best single score among its runs, so it deserved that suspicion too.

The fix is to choose the settings on one stretch of dates, then score the chosen version once on a later stretch that played no part in the choice (analysts call this nested validation). With too little data for two stretches, write predictions down before the results come back and report every finalist, not just the winner. On the household-goods work, we wrote down five predictions before a round of results came back.

How the test dates are cut. Even with no leak, the layout of the test periods moves the score. On the household-goods brand, a simple regression (trend plus media, not the full MMM) was scored on the same 32 test weeks, cut two ways. As eight four-week periods it scored +0.1751, about 18% of the rise and fall. As four eight-week periods it scored +0.1286, about 13%.

Its typical weekly miss (RMSE, root mean squared error, where 0 is a perfect forecast) grew from 857.5 to 881.4 clicks through to retailers, while the weeks' spread around their average (their standard deviation) stayed about 944 clicks. So the lower score came from bigger misses, not from a harder test. The only change was the layout: eight periods forecast up to four weeks ahead, or four periods forecast up to eight.

Ask that the test-period length match how far ahead your budget decision looks. Also ask for one score across all test weeks: a score for each short period swings too widely to mean much.

Does the model beat a calendar?

Not necessarily. A calendar forecast, one that simply repeats last year's numbers for the same week, can beat it.

A separate campaign for the same ski resort modeled revenue in six segments. There, repeating last year's numbers beat every model in that comparison in five of the six revenue segments and tied in the sixth.

A model that cannot beat that last-year forecast has not earned a budget decision. R-squared cannot show this: it only compares the model with guessing the average, and within a stretch dominated by the peak season, following the season alone can score well. That is how, the team concluded, the ski resort's Christmas period scored 0.5243.

The same yardstick can also be harsh: over a short stretch where the numbers barely move, a score worked out on that stretch alone can fall below zero even for a good forecast. So ask for the model's typical miss next to the last-year forecast's, on the same test dates and in the same units.

Then a placebo check: add a fake channel built from numbers that cannot have driven sales, and see how much credit the model gives it. In the ski-resort search, a fake channel of shuffled numbers took 32% to 51% of the credit the model reported, while the score stayed in the same range as the real models. A model that credits a fake channel can credit real ones just as wrongly.

What should you ask before you move budget?

Ask for the evidence behind each check before any channel number moves money.

Five questions for your MMM vendor or team:

Check Question to ask Failure means Decision
Later dates, with a gap (rolling origin, embargo) What were the training and test dates of every test period, and how long was the gap between them? Was every preparation step, such as rescaling, worked out on the training dates only? The model may have seen part of the answer Rerun on later dates, with a gap at least as long as the longest ad carry-over
Last-year forecast (seasonal naive) On the same dates and units, does the model beat repeating last year's numbers? Repeating last year's numbers forecasts as well as the model Do not move budget on this model
Error by test period (fold concentration) What share of the error comes from each test period? A few periods decide the score Ask what went wrong in those periods; never let anyone drop them after seeing the score
Best of many runs How many versions were tried, and how did the others score? The best score is partly luck Judge the model on the finalists' typical score, not the best one
Fake channel (placebo) How much credit did a fake channel get, and where did it rank? The model hands out credit the data cannot support Do not use per-channel numbers

The first four show whether the score reflects real forecasting skill. The fifth asks whether each channel's credit is real, which four tests for channel credit cover in full.

If the model fails the last-year comparison or the fake-channel check, treat its channel numbers as a hypothesis to test, not a budget plan. The test means switching one channel's spend on and off, or changing it in some regions only, for long enough to measure the difference, which can take months. A model whose best evidence is one good test period is still a draft, however high the score.

What does a causal modeller say about these checks?

We sent this argument to Bharath Gaddam, Global CEO of Data Poem, which builds causal models for marketing and growth planning. We asked him four questions and to disagree where he does. His answers follow as he wrote them.

When a client shows you an MMM with a strong out-of-sample score, what is the first thing you check before you believe it?

I check what the score is measuring. Out-of-sample R-squared tells you the model can predict revenue. It doesn't tell you the model knows what caused it. Those are different questions, and most MMMs, including the polynomial-regression kind, only ever answer the first. They sit on the bottom rung of Judea Pearl's ladder: association.

The ski-resort placebo result shows this well. A fake channel took 32% to 51% of the credit while the score barely moved. The authors read that as a validation problem, but I'd call it a confounding problem. Season, weather, snowfall and holidays drive both media spend and revenue. A regression that can't represent that structure will hand the credit to whatever variable correlates with the season, real channel or fake one.

So my first check is whether the model has an explicit causal structure: what drives what, what is a confounder, and what is a mediator. If it doesn't, a high score is only a well-fitted correlation.

How do you set the gap between the training and test weeks when channel effects carry over, and what do you do when the data is too short to afford it?

I agree with the article's rule that the gap must be at least as long as the longest carry-over. Leakage through adstock is real, and the household-goods DoubleML example, where the effect swung 1.5x to 2.5x from the split alone, is a useful warning.

Where I'd push back is on treating a longer gap as the fix for short data. When data is short, a purely statistical model has to relearn both the season and the media effect from very few cycles, and that is where a regression fails. A causal model can bring in structure it didn't have to learn from this one client's history, such as how carry-over and saturation behave and which relationships are stable across businesses. That is the idea behind Large Causal Models like our Fount architecture: prior causal knowledge does the work that extra weeks of data would otherwise have to do.

When data is too short even for that, I'd rather say "this can't be identified yet" than report a number. The authors' memo did exactly that, and I respect it.

Do you hold your models to a naive benchmark, like last season repeated? What happens in the room when the model loses to it?

Yes, always. A seasonal-naive baseline is the cheapest honest test there is. A model that can't beat last year's numbers hasn't earned a budget decision as a forecaster.

But I disagree with the article's conclusion that this settles the question. Forecast accuracy and causal accuracy are separate tasks. Repeating last year is a good forecast precisely because it captures everything that stays the same year to year, including the season. A budget decision, though, is a counterfactual question: what happens to revenue if I move 20% from channel A to channel B? The naive benchmark can't answer that at all, because it has no concept of an intervention. So a causal model that loses to the calendar on raw forecast error isn't necessarily useless, and one that wins isn't necessarily right about incrementality.

In the room, we separate the two conversations. We say "here is where the model loses to the calendar, and here is why," then judge the causal claims against holdouts, geo tests or other interventions. If the model fails there too, we say so.

A single holiday can carry most of a model's error. How do you handle Christmas or Black Friday without hiding them or letting them own the score?

A holiday dummy variable is a patch, and it's often the source of the problem. Christmas isn't noise to be controlled away. In a ski resort's causal graph it is a major driver of demand, and it also shapes when and how much the business spends on media. That makes it a confounder, and it belongs in the structure of the model, not in a footnote.

In practice, I do three things. I model the holiday explicitly as a driver of both demand and spend, so media doesn't inherit its lift. I report performance with and without the peak periods, in the open, so nobody can quietly drop them after seeing the score, which is the article's rule and a good one. And I check the holiday effect against a real intervention wherever one exists, such as a regional holdout or a spend pause.

I'll close with where I agree with the authors: the answer to an unreliable MMM is not more tuning, it's experiments. My addition is that experiments and causal models work best together. Experiments give you ground truth on a few channels. A causal model uses that ground truth to calibrate the rest, so you don't have to run a six-month test on every channel.