A marketing mix model (MMM) hands you a channel effect before budget moves. Precise is not causal: it is only as causal as the drivers the model saw.

We tested one brand with DoubleML (double machine learning). The model got 89 weeks of the brand's history, from late November 2024 to early August 2026, as one number per week for the whole country: clicks through to retailers. The only other things it knew about were calendar patterns: the seasons and a long-term trend.

DoubleML works in two steps. First, a forecasting method, the helper, predicts from the calendar alone how much each channel would spend and how many clicks the brand would get each week. Then DoubleML looks only at the surprises: in weeks when Google Ads spent more than the calendar predicted, did clicks also come in above prediction? The strength of that link is the channel's effect.

We ran the same data three times, changing only the helper. We tried three standard methods. A lasso is a straight-line model that keeps only the strongest patterns. A random forest averages hundreds of decision trees. Gradient boosting builds trees one after another, each fixing the last one's errors.

For Google Ads, the three helpers gave three different answers. Doubling weekly spend would add about 8 clicks a week under the lasso, remove about 5 under the random forest and add about 75 under gradient boosting. The first two are too uncertain to tell apart from zero. The third looked solid: a result that strong would turn up by chance about once in 1,400 tries if Google Ads did nothing.

Facebook shows the scale. It came out at about 1,000 to 1,460 extra clicks a week under all three helpers, and it was the only channel that kept its direction and its certainty.

Useful, not a causal estimate becoming true. DoubleML can reduce regularisation bias (the pull toward zero from penalised models) when observed controls are many and spend moves with them. It cannot adjust for a promotion, competitor move or demand shock nobody measured.

The falsifiable claim

DoubleML and a Bayesian MMM use different estimators, but observational causal identification in both rests on one claim: the controls close every back-door path, every route by which something besides spend moves both spend and outcome.

Google Ads effect under three helper methods Lasso: +8 (p 0.81) Random forest: -5 (p 0.86) Gradient boosting: +75 (p 0.00072) negative positive
Google Ads: extra weekly clicks through to retailers if its spend doubled, by helper method. A p near 1 means the result cannot be told from zero.

Is the DoubleML number a causal effect?

Only if every driver of spend and demand is a control. DoubleML fits two helper models on folds that exclude the scored week: one predicts the outcome from the controls, the other spend, then relates the two leftovers. Cross-fitting (helpers trained on other weeks) limits a helper's tendency to memorise its own row; the orthogonal score makes small helper errors matter less. Orthogonalisation is an estimator, not an instrument.

For the board: compare the demand we could not predict from known context with the spend we could not. The word known carries the identification assumption.

Outcome model, spend model, orthogonal score Outcome model Predict demand from known drivers Spend model Predict spend from those drivers Orthogonal score Relate the two leftovers Unmeasured drivers stay in both
Cross-fitting changes how the leftovers are estimated; no variation enters from outside the observed data.

Fold choice mattered as much as the learner: skip the time boundary and the number moves before you touch it. Random to contiguous blocked folds moved the same channel up to 2.49 times, same direction on a second exposure definition. Random folds let a weekly model train on neighbouring future weeks; blocked folds do not. Nothing in the library warned us.

Can an instrument replace the missing experiment?

Only a valid one; we found none. An unmeasured confounder (a driver of both spend and demand) yields only to variation that moves spend for reasons unrelated to demand: an experiment or an instrument doing that job. We screened five candidates for strength, leave-one-out stability and negative controls. None survived.

Strength is not validity. An F-statistic above 10 says the candidate predicts spend strongly enough for first-stage work, not why, nor whether it reaches the outcome. We also required leave-one-out stability (survives dropping each week) and a negative-control outcome the instrument could not plausibly affect; a strong candidate that predicts it fails.

Skip it and a CPM instrument passes. Spend equals impressions times CPM divided by one thousand, so it guarantees strength, not validity. Facebook's had an F-statistic of 470.04, still 348.69 after dropping its most influential week, and returned -294.11 with a 95% interval from -788.40 to 200.21, against all three DoubleML estimates. We discarded it because the exclusion restriction (the instrument may reach the outcome only through spend) failed by construction, not because the result was inconvenient.

Do the MMM vendors admit the same limit?

Yes, in writing. Google's Meridian calls conditional exchangeability (the controls capture everything moving both spend and outcome) the main untestable assumption behind a causal reading of an MMM regression, never perfectly met. Fair disclosure, not a defect. Meta's Robyn recommends experiments to bring causality into an MMM.

This is identification before estimation. Identification asks whether the causal quantity can be recovered from the data and assumptions at hand; estimation asks how to compute it. Cross-fitting, regularisation and Bayesian priors improve estimation under the chosen model. None supplies the missing counterfactual variation, so a better estimate of an unidentified quantity is still unidentified.

Skip the distinction and an estimate that holds under assumptions becomes a measurement on a board slide. A back-door adjustment in a Bayesian model and a cross-fitted DoubleML score both estimate after adjustment. Neither tests whether the control set was complete.

Question DoubleML helps? What answers it
Many correlated observed controls Yes Cross-fitting and orthogonalisation
Regularisation bias in helper models Yes Orthogonal score plus rate conditions
Unmeasured demand shock No Better controls, experiment or valid instrument
Spend chosen in response to expected demand No Exogenous variation and a defended design
Is the control set complete No Domain argument; observational data cannot test it alone

So what should the slide say?

A conditional one. We still use DoubleML: here it gave a learner-sensitivity map, a fold-leakage diagnosis and a robustness value (how strong an omitted confounder must be to erase the effect): 39.87% for Facebook, roughly 25% for the other channels, under the one learner where all three were significant. Under the lasso Google Ads was already at p = 0.81, so its robustness was effectively zero. Sensitivity analysis inherited the learner choice.

Given this control set, time split and helper learner, this is the association left after orthogonalisation. If the decision needs a causal number, the next dollar belongs to experimental design, not another observational estimator.