A promo that looked like $23 per customer made $6

Naive analysis: $23.03 per customer. Cross-fitted AIPW: $6.42. Planted truth: $5.87.

The question

A free-shipping promotion had 37% uptake. Did it work? The thirty-second query says yes, emphatically. But the customers who used it were the ones who open marketing emails, had been shopping for years, and already spent more than average, so the comparison measures the promo plus the fact that the best customers took it.

What I did

I built a synthetic dataset from an explicit potential-outcomes data-generating process with a planted effect, so every estimator could be scored against a known answer rather than against each other.

I ran five estimators, each against its own estimand (a naive difference in means, an OLS adjustment, propensity score matching, stabilised IPW, and cross-fitted AIPW), and kept the distinction between ATE and ATT explicit rather than treating the five numbers as five attempts at the same quantity.

Then I carried the estimate through to a business decision, netting incremental revenue against gross margin and the shipping subsidy instead of stopping at a statistically significant lift.

Results

Estimates against the planted truth.
MethodTargetEstimateTrue
Naive difference in meansdescriptive, non-causal$23.03n/a
OLS adjustedmodel-dependent adjusted coefficient$7.09$5.87†
Propensity score matchingATT$7.08$5.98
IPW (stabilised)ATE, trimmed$6.38$5.87
AIPW (cross-fitted)ATE, trimmed$6.42$5.87

† Reference benchmark, not a formal target. OLS estimates a model-dependent adjusted coefficient rather than a defined causal estimand, so $5.87 is shown for comparison rather than as the quantity this row sets out to recover.

The economics

Incremental revenue and net contribution after gross margin and shipping subsidy.
SegmentIncremental revenueNet contribution95% CI
Low spend$7.98+$1.29[+1.05, +1.51]
Mid spend$7.12+$0.25[+0.03, +0.51]
High spend$4.01−$2.09[−2.49, −1.40]

Intervals come from a joint bootstrap over both the revenue estimate and the purchase-incidence estimate.

Low and mid spend both support targeting: each is positive and each interval excludes zero. High spend destroys contribution. The shipping subsidy scales with how often people order, not with how much they spend, so high-spend customers absorb the most subsidy while showing the smallest lift. Every segment still looks profitable on incremental revenue alone, which is why a revenue-only analysis would have told you to keep subsidising exactly the wrong group.

What I'd flag in review

← Back to home