Validate on a blocked tail
Use the runner’s blocked-tail stage to evaluate a fresh aggregate model on the last configured dates. It fits earlier rows in a separate model and does not reuse the main posterior. The main run may still fit all observations, so use the holdout artefacts for the out-of-sample assessment.
validation:
enabled: true
holdout_observations: 8
include_last_observations: true
coverage_levels: [0.5, 0.8, 0.94]
posterior_predictive_random_seed: 43
sampler:
draws: 1000
tune: 1000
chains: 4
cores: 1
random_seed: 42
target_accept: 0.95
These are illustrative settings. Use MCMC and an ordinary
aggregate MMM without panel dimensions. The supported evaluation procedure here
uses no calibration. The runner rejects enabled validation with any non-empty
calibration block before loading data or creating a run directory. The same
check applies to a dry run and to direct validation-stage execution
(src/ammm/pipeline/stages/validation.py:28, src/ammm/pipeline/runner.py:174,
working tree after 7cb7f20). Remove calibration for the evaluation, or set
validation.enabled: false for a calibrated fit. --no-validation also disables
the stage during an actual run; dry runs continue to validate the source YAML
without applying stage-disable overrides. The split needs at least
max(adstock.l_max + 1, 8) earlier dates. The three coverage levels and factual
carry-in setting are fixed as shown. Set target_column explicitly in YAML.
Validation sampler entries override model sampler settings for the fresh fit;
CLI sampler overrides also apply to configured validation settings. The runner
controls progress display. --quick and --no-validation skip this stage.
Leakage and carryover
The split sorts and validates dates, extracts the target, and builds a new model on training rows with inference-data loading disabled. A separate y file must have one column aligned with X. The builder prepares target-derived holiday features using the training subset for this fit.
For the uncalibrated workflow above, prediction prepends factual training history for adstock and trims it from the scored dates. This preserves carryover from spend that occurred before the holdout, without conditioning the new posterior on held-out targets. The observed holdout media and controls are provided, so this is conditional prediction, not necessarily a forecast using only information available at the original forecast date.
Results
Under 35_holdout_validation, inspect validation_metadata.json,
holdout_predictive_report.json, holdout_predictive_summary.csv and the
retained holdout_posterior_predictive.nc. Row-level fitted, observed and
residual CSVs accompany the plots.
RMSE, MAE, CRPS, interval coverage and bias answer different predictive questions. Residuals and bias use observed minus posterior mean. NRMSE and NMAE divide by the observed holdout range and are null when it is zero. Schema 2 diagnostic policy can gate holdout NRMSE and its ratio to in-sample NRMSE; raw metrics do not carry a universal acceptance threshold.
One short tail can have noisy coverage and may miss regime changes. Repeated model selection on the same holdout weakens its role as independent evaluation. Use rolling-origin checks where appropriate, retain an independent final test when needed, and assess causal assumptions separately.
Implementation reference at 7cb7f20: src/ammm/mmm/blocked_holdout.py:76, src/ammm/pipeline/stages/validation.py:27, src/ammm/mmm/builders/schema.py:159.