We built a marketing mix model (MMM, a model that splits sales between media channels and everything else) for a household-goods brand in several versions, then checked how well each one predicted the brand's weekly clicks through to retailers on ordinary weeks it had never seen. The score is R-squared: 0 means no better than guessing the average week every time, 1 means every rise and dip predicted. The version carrying a channel made of shuffled numbers scored 0.018, under 2% of the swings. The version built from four real channels scored 0.007, under 1%. Both are close to zero, and the fake channel did not make the model worse. A model can pass as working and still not know which channels earned the sales.

That was an ablation, the standard name for this test: compare the full model with versions that remove, replace or regroup one part at a time, keeping everything else the same. If the data can see a channel's effect, removing it should hurt the forecast, and adding nonsense should not help.

A model's baseline is the revenue it expects with no media at all: trend, calendar and controls, meaning non-media factors, not a control group. When that baseline is free to bend, it can predict revenue while spreading credit almost arbitrarily among channels that move together.

A separate model, for a ski resort's revenue, tracked 13 media channels. On data it had never seen, it predicted about 40% of the rise and fall in revenue: an R-squared of 0.4028. With all 13 channels removed, it scored the same, 0.4057, and a version with the controls alone scored 0.4027. So the 40% came from trend, season and weather, not from media. The money shows it too: over 693 days the model put about $39.2 million of revenue down to the baseline and $3,351 to all 13 channels together, less than one hundredth of a percent. It still gave $250 of that to two channels with zero spend. Credit for a channel with zero spend is a red flag on its own.

Key Takeaways

Trust a channel's credit only if it survives four attacks: remove media, add a placebo, invert the prior, regroup channels.

  • A ski-resort model forecast just as well with all media removed, yet credited $250 to two channels with zero spend.
  • On a household-goods brand, the version carrying a channel of shuffled numbers scored higher out-of-sample than one built from four real channels.
  • Automatic clean-up rules can merge or delete a channel on their own: one flagship paid-search channel was deleted because a protection list missed its name.
  • Ask your MMM vendor or team which of the four tests each channel number survived.
Ski-resort model: three near-identical out-of-sample scores Out-of-sample R-squared Controls only 0.4027 Champion, with media 0.4028 All media removed 0.4057
Ski-resort revenue model: removing media did not hurt the forecast, so fit cannot vouch for media credit.

How do you know a channel's credit is real?

Attack it four ways, and have whoever runs the model write down first what a failure means.

  1. Remove media entirely: does media improve the forecast beyond trend, seasonality and controls?
  2. Add a placebo channel: real spend shuffled, or rolled in time so it lines up with weeks it cannot have caused. Does the model credit it anyway?
  3. Invert the prior, the starting belief we hand the model: give it the opposite assumption. The further the credit moves, the more of the answer came from us, not the data.
  4. Regroup channels: do individual estimates survive being merged with the channels they move with?
Four ablations in sequence Remove media Drop all media Add a placebo Shuffled fake spend Invert the prior Assume the opposite Regroup channels Merge channels
Run in this order.

It is tempting to read each failure as its own scandal. They are one sequence. If removing media shows that the baseline already predicts sales, a fake channel winning credit is the expected next result, not a second discovery: the model cannot tell the channels apart. If inverting the prior then shifts the split while the forecast barely moves, the three results say one thing: the sales data decides the baseline, our assumptions decide the split.

Can automatic clean-up rules change which channel gets credit?

Yes: merging or deleting overlapping channels can erase distinctions the business cares about.

Correlation, written r, measures how closely two channels move together. It runs from -1 to 1: near 1 they rise and fall almost in lockstep, near 0 they barely move together, below 0 one rises as the other falls. When two channels sit above 0.90 or below -0.90, the data cannot tell their effects apart, and any split of credit between them is a guess. So our engine can merge such a pair into one channel.

That is a forced trade, not a bonus: one number the data supports replaces two it cannot, and the business loses the separate answer. In one retail run, brand search and brand shopping merged at 0.954, then non-brand shopping joined them at 0.947, and the client could no longer ask what non-brand shopping alone was worth.

The dates you measure also decide whether a pair belongs together. Two programmatic ad platforms scored 0.980, 0.966 and 0.959 across three different date ranges, so merging them was right. A social channel and brand search scored 0.90 within one year but -0.190 over the full date range, so they did not belong together. If you need two merged channels apart, change the spend of one of them on its own for a while, for example in some regions and not others, so the data has something to separate them by.

Deleting channels carries the same risk. In our engine a variance inflation factor (VIF, how much a channel duplicates the others) above 10 triggers removal, unless the channel is on an exact-name protection list checked first. In a real run, a flagship paid-search channel at 10.04, barely over the line, was deleted: the protection rule missed its name. The model did what the code asked. The code did not represent the business rule. Ask which channels the model merged or deleted on its own.

What should the report say when a claim fails?

It should refuse the claim. Ablation does not prove causality. What it shows is whether a claim is sturdy. A claim that still holds when we change the model's settings in sensible ways is worth using; one that appears in only one version of the model is not.

Refusal cuts both ways: a test result can be refused too. We once withdrew two placebo results, at 32% and 41% of contribution, because the fake channel had been let off an assumption we imposed on the real ones and neither fit had settled on a stable answer. A fake channel has to be held to the same standard as a real one, or the result decides nothing.

On a portrait-studio account, placebo channels rolled by 183 and 274 days ranked last among eleven channels. A 91-day roll still captured 8.0% of contribution. That one is the finding: a channel that cannot have caused those weeks' sales earned credit, so on this account a real channel at or below 8.0% has not shown it is more than noise.

We then ran sixteen model setups. One channel, 58.7% of spend, stayed inside the efficiency band (credit-per-dollar limits fixed in advance) in all sixteen. The other 41.3% of spend did not earn a channel-level claim. A useful report names the few channel numbers it will stand behind, and which it will not.

Four questions for your MMM vendor or team:

Ablation Question to ask Failure means Decision
Remove media Does the forecast get worse without media? Attribution is not required for fit Do not treat fit as proof of attribution
Add a placebo Can a fake channel win credit? Noise wins credit from the baseline or our priors Do not use per-channel numbers
Invert the prior Do channel values change if we assume the opposite? The data are too weak to split credit Report sensitivity, not certainty
Regroup channels Do channel values survive regrouping? The data cannot tell these channels apart Report the bundle

Meridian, Google's open-source MMM, documents that a model with 99% out-of-sample R-squared can still be poor for causal inference. We would add: a model that predicts as well after you remove the thing it claims to measure has failed a simpler test. Before you move budget on a channel's number, ask your MMM vendor or team which of the four ablations it survived, and apply the table above to any it failed.