Should everyone keep seeing this campaign?
A decision memo for a growth team: is the ad campaign working, how sure are we, and if we stop showing it to everyone, who should still get it?
The short version
- The campaign works. Showing it lifts the conversion rate by 0.100pp (95% CI 0.092pp to 0.107pp), a relative lift of 49% over users who were held out. That is about 1.0 extra conversions per 1,000 users targeted.
- The effect is concentrated in a small group. Rank users with a model trained on other users. The top 10% then account for 82% of the extra conversions, and the top 20% for 88%. Across five held-out folds those shares range from 79% to 88% and from 85% to 91%.
- Recommendation: stop showing the campaign to everyone; show it to the model's top slice. Blanket targeting only beats targeting if reaching a user costs less than about 15% of the average extra value per user, a very low bar. Suppose instead that a targeted user costs half the average extra value. Then the best cut-off is the top 16%, and it nets 1.6× what blanket targeting does. At cost equal to the average value, blanket targeting only breaks even, while the top 10% still makes money.
- One caution before acting: this public sample is not a perfectly clean randomization (§1.2). The headline numbers use an estimator that corrects for it. A naive treated-vs-control comparison would report a 59% conversion lift instead of 49%.
1 · Experiment readout
The data are 13,979,592 users from Criteo incrementality tests. Treated users were eligible to see the campaign's ads, and control users were held out. Two outcomes: whether the user visited the advertiser's site, and whether they converted. Twelve anonymized features were recorded before treatment. Everything below is computed on all rows, with DuckDB aggregates over a parquet file (peak memory under 2 GB).
1.1 Was the split what the design said it was?
The test was designed as an 85/15 split. The observed treated share is 0.8500001 (11,882,655 treated, 2,096,937 control). A chi-square sample-ratio-mismatch test against 0.85 gives χ² = 1.8e-06, p = 0.999. There is no sample ratio mismatch. (Testing against a naive 50/50 expectation would "fail" with a p-value that underflows to zero, which is why the test has to be run against the design ratio.)
1.2 Covariate balance: the check that fails
Randomization should make the two groups look alike before treatment. With groups this large, chance alone produces standardized mean differences (SMD) with a standard deviation of about 0.00075. The observed SMDs reach 0.049. All 12 of 12 features differ significantly after Holm correction. A gradient-boosted model can predict treatment from the features on held-out rows with AUC 0.5075. Twenty label permutations give a maximum AUC of 0.5003.
This fits the dataset's own documentation. The release was assembled from several incrementality tests and, for privacy, was "sub-sampled non-uniformly so that the original incrementality level cannot be deduced." Either could tie assignment to the features. The practical consequence is that a raw treated-minus-control difference is not a clean causal estimate here. The analysis therefore treats assignment as random conditional on the features, with an estimated propensity e(x). Its middle 90% spans 0.832–0.877, so overlap is excellent. The headline estimator is augmented inverse-propensity weighting (AIPW, doubly robust), with nuisance models cross-fitted on two halves of the data.
1.3 Average treatment effect
CI 0.092pp to 0.107pp
CI 43.6% to 53.6%
CI 0.72pp to 0.78pp
CI 17.8% to 19.3%
| Estimate | Visit | 95% CI | Conversion | 95% CI |
|---|---|---|---|---|
| Control rate (raw) | 3.820% | 0.194% | ||
| Treated rate (raw) | 4.854% | 0.309% | ||
| Difference in means, analytic | 1.034pp | 1.006pp – 1.063pp | 0.1152pp | 0.1085pp – 0.1219pp |
| Difference in means, bootstrap (10,000 resamples) | 1.005pp – 1.063pp | 0.1083pp – 0.1217pp | ||
| Relative lift, naive (delta method) | 27.1% | 26.2% – 28.0% | 59.4% | 54.4% – 64.7% |
| AIPW (doubly robust) | 0.751pp | 0.725pp – 0.777pp | 0.0996pp | 0.0923pp – 0.1068pp |
| Relative lift, AIPW | 18.6% | 17.8% – 19.3% | 48.5% | 43.6% – 53.6% |
| Effect on users actually exposed (AIPW ÷ exposure rate) | 20.84pp | 20.11pp – 21.57pp | 2.76pp | 2.56pp – 2.96pp |
The analytic and bootstrap intervals agree to the displayed precision. The bootstrap resamples rows with replacement; for a binary outcome that is exactly a binomial draw per arm, so it needs no pass over the data. Only 3.6% of treated users were actually shown an ad, and control users never were. Dividing the intent-to-treat effect by that rate (a Wald / instrumental-variable ratio) estimates the effect on the users who saw it. That assumes being eligible but unexposed has no effect of its own.
1.4 Tighter intervals from pre-treatment data
The twelve features predict the outcomes well, so adjusting for them removes noise. CUPED uses one pooled linear coefficient. Lin's estimator lets the slope differ by arm. CUPAC uses a cross-fitted gradient-boosted prediction as the single covariate.
| Estimator | Visit effect | Variance vs raw | Conversion effect | Variance vs raw |
|---|---|---|---|---|
| difference in means | 1.034pp | 100% | 0.1152pp | 100% |
| CUPED, 12 linear features | 0.698pp | 75% | 0.0916pp | 89% |
| Lin interacted regression | 0.773pp | 74% | 0.1002pp | 89% |
| CUPAC, cross-fitted GBM score | 0.638pp | 70% | 0.0939pp | 89% |
| AIPW, cross-fitted GBM | 0.751pp | 85% | 0.0996pp | 116% |
CUPAC cuts the variance of the visit estimate to 70% of the raw estimator's. That is the same precision as running the test on 1.44× as many users. For the rare conversion outcome the gain is smaller (89%). AIPW is less precise than the pure regression adjustments on conversion (116% of raw variance), because inverse-propensity terms add noise when outcomes are rare. It is still the headline. With assignment tied to the features, it is the only estimator here that corrects through both a propensity model and an outcome model, so it stays consistent if either one is right. The adjusted point estimates disagree with each other by more than their own intervals (visit: 0.64pp to 0.77pp). That is model dependence left over after the imbalance, so the quoted CIs understate the real uncertainty.
1.5 Who responds more? Segment effects
Each feature was cut at its deciles. Several features have one dominant value, so only 9 of 12 split into two or more segments, giving 80 segment-by-outcome effects. Each segment's effect is the mean AIPW score of its users. Correcting for 80 comparisons, 76 remain significant under Holm (family-wise) and 76 under Benjamini–Hochberg (false discovery rate). 0 of them are negative: no segment was measurably harmed. A heterogeneity test per feature (Cochran's Q across its segments, Holm-corrected over 18 tests) rejects "same effect in every segment" for 18 of 18.
2 · Power and design
2.1 How big does the next test need to be?
The control conversion rate is 0.194%. At rates that low, small relative lifts need very large samples. With 14.0M users at 85/15, the smallest detectable conversion lift (80% power, two-sided α = 0.05) is 0.0092pp, or 4.8% relative. The 85/15 allocation costs precision. Its variance is 1.96× that of a 50/50 split of the same size, so a test that reserves few users for control needs about twice as many users in total.
| Relative lift to detect | Users, 50/50 | Users, 85/15 |
|---|---|---|
| 2.0% | 40,833,518 | 79,511,898 |
| 3.0% | 18,237,892 | 35,391,364 |
| 5.0% | 6,630,194 | 12,778,863 |
| 7.5% | 2,982,612 | 5,700,589 |
| 10.0% | 1,697,889 | 3,218,445 |
| 15.0% | 772,543 | 1,440,965 |
| 20.0% | 444,637 | 816,473 |
| 30.0% | 206,575 | 368,147 |
| 50.0% | 80,813 | 136,325 |
An effect as large as the one measured here (49% relative) needs only 144,581 users at 85/15. With CUPAC's variance ratio that falls to about 128,877. The next useful test is a targeted one (§3). The effects it has to detect are smaller, so the table above is the one to plan with.
2.2 Why you cannot stop the test the first time it looks significant
This section simulates 100,000 experiments per setting in which the treatment does nothing, and checks the p-value after each batch of data. Stopping at the first p < 0.05 inflates the false-positive rate well past the promised 5%:
There are two standard fixes. The table below compares them on 20 equally spaced looks, for a design with 80% power at the planned sample size. Alpha spending (Lan–DeMets with an O'Brien–Fleming-type spending function) fixes the looks in advance and uses very strict thresholds early. mSPRT (a mixture sequential probability ratio test, as used in always-valid p-values) allows checking at any time, with no plan.
| Method | False positive rate | Power | Avg. sample used (if effect is real) |
|---|---|---|---|
| Fixed horizon (look once) | 4.9% | 79.9% | 100% |
| Naive peeking (z > 1.96 at any look) | 24.7% | 88.2% | 45% |
| O'Brien-Fleming alpha spending | 4.9% | 77.3% | 74% |
| mSPRT (always-valid) | 1.1% | 49.0% | 82% |
Alpha spending keeps the error rate at 5%, gives up little power, and stops early on average when the effect is real. mSPRT is conservative at the planned horizon (its guarantee holds for unlimited looks). Its cost is sample size: if it may run to 2.0× the planned sample, power reaches 87%, while the average sample used is only 1.11× the plan. Use alpha spending when the review schedule is fixed. Use mSPRT when stakeholders will look at a live dashboard.
Simulation details
The z-statistic path is a Gaussian random walk (exact for known-variance difference in means; an excellent approximation at thousands of conversions per look). Alpha-spending boundaries were calibrated on an independent set of 400,000 null paths. The mSPRT uses a normal mixing prior with variance τ² = 0.392 per look, set to the planned effect size, and rejects when the mixture likelihood ratio exceeds 1/α.
3 · Who should see the campaign? Uplift modeling
An uplift model predicts, for each user, how much the campaign changes their chance of converting. Four standard learners were compared, all using LightGBM as the base model:
- S-learner: one model with treatment as a feature; uplift = prediction with minus without.
- T-learner: separate models for treated and control users.
- X-learner: T-learner models impute each user's individual effect; a second stage learns those; the two are blended by the propensity.
- Transformed outcome: regress Y·(T/e − (1−T)/(1−e)), whose expectation is the individual effect.
Honest evaluation. Users were split into 5 disjoint folds. For each fold, every learner was trained on 1,500,000 users drawn from the other four and scored on all ~2.8M users of the held-out fold. A slice's uplift is measured with the held-out users' AIPW scores, so it is unbiased even with the imbalance from §1.2. Curves, areas and policy values below never touch training rows, and the spread is across the five folds.
| Learner | Conversion Qini coef. | Conversion AUUC ×10⁴ | Visit Qini coef. | Visit AUUC ×10³ |
|---|---|---|---|---|
| T-learner | 0.513 ± 0.060 | 7.54 ± 0.97 | 0.631 ± 0.036 | 6.12 ± 0.23 |
| X-learner | 0.534 ± 0.071 | 7.64 ± 0.98 | 0.693 ± 0.012 | 6.36 ± 0.23 |
| S-learner (best) | 0.772 ± 0.053 | 8.82 ± 1.09 | 0.837 ± 0.022 | 6.90 ± 0.24 |
| Transformed outcome | 0.329 ± 0.107 | 6.61 ± 0.92 | 0.571 ± 0.025 | 5.89 ± 0.16 |
| Random | 0.008 ± 0.061 | 5.03 ± 0.80 | 0.009 ± 0.030 | 3.79 ± 0.24 |
AUUC is the area under the uplift curve (per-user scale). The Qini coefficient is the model's area above the random-targeting diagonal, divided by the area under that diagonal: 0 means no better than random, 1 means twice the area. Values are mean ± standard deviation across the five folds. On conversion, the S-learner has AUUC 8.82×10⁻⁴ against 5.03×10⁻⁴ for a random ranking (1.75×). It is the simplest learner and it wins by a margin larger than the fold-to-fold spread. That fits a known pattern: when the effect is small next to the baseline, fitting the two arms separately (T-, X-learner) or fitting a high-variance transformed target mostly adds noise. "Best" is picked on the same held-out folds, so its lead over the runner-up is slightly optimistic. The gap between any real learner and random is not.
Are the predictions calibrated?
What is it worth? Policy value under a cost assumption
The dataset contains no prices, so costs are stated as a ratio. Let r be the cost of targeting one user divided by the value of one conversion. Targeting the top φ of users is worth V·[uplift(φ) − r·φ] per user in the population, where uplift(φ) is the held-out curve above. The best φ depends only on r:
| Scenario: cost per targeted user | Best share to target | Net value, best policy | Net value, target everyone |
|---|---|---|---|
| 0.02× avg. incremental value per user | 100% (100%–100%) | +9.76 | +9.76 |
| 0.1× | 100% (15%–100%) | +9.00 | +8.96 |
| 0.5× | 16% (14%–19%) | +7.96 | +4.98 |
| 1.0× (blanket breaks even) | 10% (7%–13%) | +7.26 | +0.00 |
| 1.5× (blanket loses money) | 7% (7%–11%) | +6.83 | -4.98 |
Net value is in conversions' worth per 10,000 users in the population (multiply by the value of one conversion to get money), averaged over the five held-out folds. The best-share ranges are across folds.
4 · Recommendation and limitations
For the growth team, in plain terms: the campaign causes real extra conversions, but about four in five of them come from roughly one user in ten, and a model can find those users in advance. For everyone else the campaign does very little. Unless showing the ad is nearly free, stop showing it to everyone and show it to the top-ranked 10–20%. That keeps most of the gain for a fraction of the cost. Do not read the model's bottom-decile "negative" predictions as users to protect from the ad. §3 shows those predictions are wrong. The next step is to test that directly: a new randomized test comparing "target the model's top slice" against "target everyone", sized with §2 and monitored with alpha spending.
Limitations.
- The effect sizes are not Criteo's real ones. The release was sub-sampled on purpose so that the true incrementality cannot be recovered. The directions, the ranking of users and the methods carry over. The exact percentages do not.
- The sample is not cleanly randomized (§1.2). The adjusted estimates assume that the anonymized features capture whatever tied assignment to users. They remove most of the measured imbalance but not all of it, and different adjustment models disagree by more than their stated CIs. Treat the headline CI as a lower bound on the true uncertainty.
- Anonymized features. The features are random projections of the originals. Segments can be targeted but not explained: we cannot say who the high-uplift users are in business terms.
- Ad context and external validity. These are historical tests for particular advertisers and placements. Uplift models age as creatives, auctions and audiences change, so a targeting policy needs its own holdout to stay honest.
- Costs are assumed. No prices are in the data; the policy results are stated as cost-to-value ratios so they can be re-read with real numbers.
Reproduce
Code, tests and this page are generated from one repository: github.com/tachyurgy/uplift-readout. make data downloads and checksums the file and converts it to parquet. make analysis runs the nuisance models, readout, power simulations and uplift models (about 30 minutes on a laptop, under 2 GB of memory). make site renders this page from the result files; no number on it is typed by hand.