Predicted incrementality with PIE

PIEModel predicts the incrementality of a campaign that has not been tested, using a BART regression fitted on campaigns whose incrementality an experiment measured. It follows Gordon, Moakler and Zettelmeyer (2026), NBER Working Paper 35044, but uses Bayesian Additive Regression Trees, which give a posterior for each prediction where the paper’s random forest gives a point estimate (src/ammm/pie/model.py:80). The module is alpha, so its API and defaults may change. Install it with the pie extra.

When PIE is appropriate

PIE suits an organisation that has run a corpus of comparable experiments (geo tests, conversion-lift or ghost-ad holdouts) and needs an estimate for campaigns that cannot each be tested. Its predictions transfer the experiments to new campaigns, so they are valid only if the corpus represents the campaigns being predicted and the recorded features explain the variation in incrementality. A few dozen experiments spanning the target campaigns is a practical minimum; with fewer, the posterior intervals are wide and the held-out check below has little power.

Data contract

Each row of X is one experiment and y holds its measured incrementality, in units you choose and apply consistently, such as incremental conversions per unit of spend. The model reads only the columns it is configured to use and ignores the others (src/ammm/pie/model.py:302).

ArgumentRole
pre_determined_featuresColumns known before launch, such as objective, vertical, planned budget or audience type. Always used.
post_determined_featuresColumns observed only after launch, such as exposure rate, click-through rate or last-click conversions. Used only when use_post_determined_features=True.
use_post_determined_featuresTrue by default. Set it to False for predictions made before launch; the model then neither trains on nor requires the post-launch columns.
standard_error_columnOptional column of per-experiment standard errors in the units of y.
target_columnName of the target in the PyMC graph and saved inference data. It does not select a column from X.

Non-numeric columns are label-encoded and split one level at a time. Categories absent from training raise an error at prediction time. Missing or infinite feature values raise an error at fit and prediction time because PIEModel does not impute; drop or fill those rows and record how (src/ammm/pie/model.py:279). A numeric prediction input outside the training range raises a UserWarning, because BART holds its prediction constant beyond the edge of the corpus (src/ammm/pie/model.py:465).

Choose the feature boundary from the decision time. Post-launch features often carry most of the predictive signal, but a prediction made before launch cannot use them. The paper also splits the sample within each campaign so that post-launch features are not computed from the same outcomes as the measured lift; AMMM4 does not implement that split, so a post-launch feature derived from the experiment’s own outcome can overstate accuracy.

Measurement error

Supplying standard_error_column stops the model from treating every experiment as exact. The likelihood for experiment i becomes Normal(f(x_i), sqrt(sigma**2 + se_i**2)), which marginalises the latent true incrementality, so a precise experiment carries more weight than a noisy one (src/ammm/pie/model.py:437). Predictions set the standard error to zero, so their draws describe true incrementality rather than a new noisy measurement (src/ammm/pie/model.py:497). In a synthetic check where half the corpus measured 1.0 with standard error 0.01 and half measured 0.0 with standard error 5.0, the mean prediction was 0.51 without the column and 0.96 with it (tests/pie/test_model.py, test_standard_errors_downweight_noisy_experiments).

Fit a pre-launch model

The following synthetic example fits a pre-launch model with standard errors.

import numpy as np
import pandas as pd
from ammm.pie import PIEModel

rng = np.random.default_rng(1)
n = 120
corpus = pd.DataFrame({
    "objective": rng.choice(["conversions", "traffic"], size=n),
    "planned_budget": rng.uniform(5_000, 100_000, size=n),
    "audience_size": rng.uniform(1e5, 5e6, size=n),
    "lift_se": rng.uniform(0.02, 0.2, size=n),
})
true_lift = (
    0.4 + 0.3 * (corpus["objective"] == "conversions")
    + corpus["planned_budget"] / 2e5
)
lift = pd.Series(rng.normal(true_lift, corpus["lift_se"]), name="lift")

def make_model():
    return PIEModel(
        pre_determined_features=["objective", "planned_budget", "audience_size"],
        post_determined_features=[],
        use_post_determined_features=False,
        standard_error_column="lift_se",
        target_column="lift",
        model_config={"bart": {"m": 50, "alpha": 0.95, "beta": 2.0}},
    )

pie = make_model()
pie.fit(corpus, lift, draws=500, tune=500, chains=2, random_seed=42)

The default configuration uses 200 trees; this example uses 50 to keep the run short. Check the usual sampler diagnostics before reading the predictions. PGBART does not mix in the way NUTS does, so an R-hat above 1.01 on the bart variable is common; run more draws and chains and compare predictions between chains rather than relying on R-hat alone.

Predict and read the uncertainty

Prediction does not need the standard-error column, and it replaces the model’s feature data, so run the diagnostics in the next section first.

new_campaigns = corpus.drop(columns="lift_se").head(5)
draws = pie.predict_posterior(new_campaigns, extend_idata=False)
summary = draws.quantile([0.05, 0.5, 0.95], dim="sample")

The 5 to 95 percent interval from predict_posterior reflects the uncertainty in the BART function given the corpus. It excludes three sources of error: whether the corpus represents the predicted campaign, whether the features capture what drives incrementality, and any bias in the experiments themselves. Treat a narrow interval for a campaign unlike the corpus as a symptom of extrapolation rather than as evidence of precision.

Feature diagnostics with pymc-bart

pymc-bart already provides interpretation tools that accept a fitted PIEModel. Run them immediately after fit and before any prediction. The feature matrix is the encoded training data, and the labels follow the model’s feature order.

import pymc_bart as pmb

X_train = pie.model["X"].get_value()
labels = list(pie.model.coords["feature"])

vi = pmb.compute_variable_importance(
    pie.idata, pie.model["bart"], X_train,
    model=pie.model, method="VI", random_seed=42,
)
pmb.plot_variable_importance(vi, labels=labels)
pmb.plot_pdp(
    pie.model["bart"], X_train, xs_interval="quantiles",
    var_idx=[labels.index("planned_budget")], random_seed=42,
)
pmb.plot_variable_inclusion(pie.idata, X_train, labels=labels)

Variable importance shows how much predictive accuracy is lost when each feature is removed, whereas variable inclusion counts how often trees split on a feature, which is cheaper but less reliable when features are correlated. Partial dependence plots show the marginal predicted response to one feature. All three describe the fitted predictor; none shows that a feature causes a change in incrementality, and partial dependence for a categorical feature is on its encoded integer scale.

Held-out coverage check

The most useful acceptance check is whether held-out experiments fall inside their predicted intervals at the stated rate. The following K-fold loop refits on four fifths of the corpus, predicts the remaining fifth, adds back each held-out experiment’s measurement error and records whether the measured lift falls inside the 90 percent interval.

from sklearn.model_selection import KFold

covered = []
for train_idx, test_idx in KFold(5, shuffle=True, random_state=0).split(corpus):
    fold = make_model()
    fold.fit(corpus.iloc[train_idx], lift.iloc[train_idx],
             draws=300, tune=300, chains=2, random_seed=42)
    held_out = corpus.iloc[test_idx]
    tau = fold.predict_posterior(held_out, extend_idata=False).values
    se = held_out["lift_se"].to_numpy()[:, None]
    measured_draws = tau + rng.normal(0.0, se, size=tau.shape)
    low, high = np.quantile(measured_draws, [0.05, 0.95], axis=1)
    measured = lift.iloc[test_idx].to_numpy()
    covered.append((measured >= low) & (measured <= high))

coverage = np.concatenate(covered).mean()

On this synthetic corpus the coverage was 0.93 against a nominal 0.90. Coverage well below nominal on a real corpus means the intervals are too narrow, which is typically caused by missing features, unrecorded standard errors or heterogeneity between experiment types. Random folds test interpolation within the corpus. To test transfer to a new advertiser, vertical or market, hold out whole groups with GroupKFold, because the paper’s cold-start diagnostics are not implemented.

Limits of PIE predictions

A PIE prediction is a model estimate, not experimental evidence, and it is not yet a suitable lift-test calibration input for an MMM. Lift-test calibration needs a spend change, a response change and its standard error (x, delta_x, delta_y and sigma) from a study of the modelled channel (src/ammm/mmm/lift_test.py:384). A PIE prediction is learned from experiments that may already calibrate the same MMM, so adding it would count that evidence twice, and its interval omits the transfer error described above. Keep measured experiments as the calibration input and report PIE predictions alongside the MMM.

AMMM4 does not implement three parts of the paper:

  • within-campaign sample splitting for post-launch features (paper section 4.2);
  • extrapolation and cold-start diagnostics across advertiser segments (paper section 5.3);
  • the decision framework comparing go and no-go choices with experiment-based decisions (paper section 6).

For the other optional integrations, see PIE and optional integrations.

Implementation reference: src/ammm/pie/model.py:80, src/ammm/mmm/lift_test.py:384.