Validate on a blocked tail

Use the runner’s blocked-tail stage to evaluate a fresh aggregate model on the last configured dates. It fits earlier rows in a separate model and does not reuse the main posterior. The main run may still fit all observations, so use the holdout artefacts for the out-of-sample assessment.

validation:
  enabled: true
  holdout_observations: 8
  include_last_observations: true
  coverage_levels: [0.5, 0.8, 0.94]
  posterior_predictive_random_seed: 43
  sampler:
    draws: 1000
    tune: 1000
    chains: 4
    cores: 1
    random_seed: 42
    target_accept: 0.95

These are illustrative settings. Use MCMC and an ordinary aggregate MMM without panel dimensions. The supported evaluation procedure here uses no calibration. The runner rejects enabled validation with any non-empty calibration block before loading data or creating a run directory. The same check applies to a dry run and to direct validation-stage execution (src/ammm/pipeline/stages/validation.py:28, src/ammm/pipeline/runner.py:174, working tree after 7cb7f20). Remove calibration for the evaluation, or set validation.enabled: false for a calibrated fit. --no-validation also disables the stage during an actual run; dry runs continue to validate the source YAML without applying stage-disable overrides. The split needs at least max(adstock.l_max + 1, 8) earlier dates. The three coverage levels and factual carry-in setting are fixed as shown. Set target_column explicitly in YAML.

Validation sampler entries override model sampler settings for the fresh fit; CLI sampler overrides also apply to configured validation settings. The runner controls progress display. --quick and --no-validation skip this stage.

Leakage and carryover

The split sorts and validates dates, extracts the target, and builds a new model on training rows with inference-data loading disabled. A separate y file must have one column aligned with X. The builder prepares target-derived holiday features using the training subset for this fit.

For the uncalibrated workflow above, prediction prepends factual training history for adstock and trims it from the scored dates. This preserves carryover from spend that occurred before the holdout, without conditioning the new posterior on held-out targets. The observed holdout media and controls are provided, so this is conditional prediction, not necessarily a forecast using only information available at the original forecast date.

Results

Under 35_holdout_validation, inspect validation_metadata.json, holdout_predictive_report.json, holdout_predictive_summary.csv and the retained holdout_posterior_predictive.nc. Row-level fitted, observed and residual CSVs accompany the plots.

RMSE, MAE, CRPS, interval coverage and bias answer different predictive questions. Residuals and bias use observed minus posterior mean. NRMSE and NMAE divide by the observed holdout range and are null when it is zero. Schema 2 diagnostic policy can gate holdout NRMSE and its ratio to in-sample NRMSE; raw metrics do not carry a universal acceptance threshold.

One short tail can have noisy coverage and may miss regime changes. Repeated model selection on the same holdout weakens its role as independent evaluation. Use rolling-origin checks where appropriate, retain an independent final test when needed, and assess causal assumptions separately.

Implementation reference at 7cb7f20: src/ammm/mmm/blocked_holdout.py:76, src/ammm/pipeline/stages/validation.py:27, src/ammm/mmm/builders/schema.py:159.