How do you validate a marketing forecast?
A forecast is validated on periods not used for fitting, with a realistic forecast origin, a naive benchmark and a metric aligned with error cost.
Scientific editorial team: Marketing Science Center
Direct answer
Compare models for a declared horizon and error cost.
A forecast is validated on periods not used for fitting, with a realistic forecast origin, a naive benchmark and a metric aligned with error cost.
01 · PROTOCOL
Operational summary
A weekly forecasting model is refitted at thirteen successive origins, then judged on 52 weeks it has never seen. Its scaled absolute error is 0.420792, against 1.039025 for a seasonal naive forecast: it clearly beats the benchmark. Its 80% intervals contain the outcome 0.750000 of the time, inside the declared tolerance of 0.10 but on the narrow side. Both checks pass. Verdict: FORECAST_READABLE_FOR_PLANNING.
02 · PROTOCOL
Concrete marketing situation
A team must commit stock and media budgets four weeks ahead. It has a model that forecasts weekly revenue and that, on the history, fits the data remarkably well. The question to ask before using it is simple: when has this model ever faced weeks it did not know, and what happened? Without that answer, the observed goodness of fit says nothing about the intended use.
03 · PROTOCOL
Scientific question
For this weekly series, this declared model, these thirteen declared origins and this four-week horizon, is the model’s error on weeks not used to fit it lower than that of a seasonal naive forecast brought to the same scale? And do its 80% intervals contain the outcome as often as they claim? The two questions are distinct and each gets its own answer: an accurate point forecast can come with a false uncertainty.
04 · PROTOCOL
Why the simple approach can fail
Judging a model on the weeks used to fit it always flatters: the more parameters one adds, the better the result looks, without any forecasting ability having been demonstrated. Splitting the history once into training and test is not enough either, since the score then depends on the accident of a single boundary. And reporting an error in currency units with no benchmark leaves nobody able to judge: a mean error of 55 units is excellent or catastrophic depending on the series.
05 · PROTOCOL
Method intuition
The method reproduces real use. One steps back to a past date, uses only what was known that day, forecasts the following four weeks, then moves four weeks forward and starts again. Thirteen times. Each error is then divided by a reference quantity computed on the past available at that date alone: the error a seasonal naive forecast would have made. A value below one means the model did better than that simple rule. In parallel, one counts how often the announced interval actually contained the outcome.
06 · PROTOCOL
Required data
Frozen before any reading: the gapless week numbering, the weekly revenue, the model form (constant, linear trend and two annual harmonics), the seasonal period of 52 weeks, the thirteen origins from 104 to 152 in steps of 4, the horizon of 4 weeks, the seasonal naive forecast as benchmark, the scaling quantity, the nominal level of 0.80, the quantile 1.281552 and the two thresholds. The sealed file holds 156 weeks. The case is synthetic and declared as such.
07 · PROTOCOL
Formal model and symbols
At each origin, the model fits by least squares a constant, a linear trend and two annual harmonics — six parameters — on the weeks before the origin only. The scaled absolute error divides each out-of-sample error by the mean of the 52-week absolute differences observed in that same window. The interval has a half-width of 1.281552 times the residual standard deviation of the fit, that quantile being the 0.90 quantile of the standard normal law. The interval score adds to the width a penalty proportional to any miss, divided by the accepted risk.
08 · PROTOCOL
Declared calculation
The declared computation runs seven steps: validate the schema and refuse any non-conforming file or one too short for the validation design; at each origin, build the design matrix on the earlier weeks only; solve the normal equations by Gaussian elimination with partial pivoting, refuse a singular window and compute the residual standard deviation with the six parameters deducted from the degrees of freedom; compute the seasonal scale on the same window and refuse a zero scale; for each of the four horizons, record the absolute error, the scaled error, the naive one, whether the interval contains the outcome, the width and the score; average over the 52 forecasts and then separately by horizon; apply the two declared thresholds, then the verdict rule.
09 · PROTOCOL
End-to-end numeric example
Synthetic illustration — teaching values, not observed
On the sealed file: 156 weeks, 13 origins from week 104 to week 152, horizon 4, hence 52 out-of-sample forecasts. Model scaled absolute error 0.420792, against 1.039025 for the seasonal naive forecast; mean absolute error 54.912955 currency units; accuracy check passed. By horizon: 0.418133, 0.520418, 0.279945 and 0.464672. Nominal level 0.80, observed coverage 0.750000, mean width 168.860643 and mean interval score 243.378776; calibration check passed. Verdict: FORECAST_READABLE_FOR_PLANNING.
10 · PROTOCOL
Validity assumptions
Validation assumes that the mechanism which produced past weeks is still the one that will produce future weeks, that the seasonal period really is 52 weeks, and that historical values have not been revised after the fact. It also assumes, for the intervals, that the residuals behave like normal noise with constant variance. Three things are not testable here: the absence of a future regime change, the effect of a channel or a price that never varied in the history, and the quality of the data at the moment the forecast will actually be produced.
11 · PROTOCOL
Diagnostics and uncertainty
Two caveats deserve to be read before the verdict. The observed coverage is 0.750000 for an announced level of 0.80: the gap stays inside the declared tolerance, but the intervals are on the narrow side, which is exactly what one expects from intervals built on residual spread alone, ignoring both parameter uncertainty and its growth with the horizon. And the per-horizon errors — 0.418133, 0.520418, 0.279945, 0.464672 — do not increase steadily with the horizon: with thirteen forecasts per horizon that pattern is noise, and no threshold is attached to them.
12 · PROTOCOL
Robustness and alternatives
Alternatives declared before results: replace the normal intervals by intervals built on the empirical quantiles of the out-of-sample errors, one set per horizon; add a simple random walk to the benchmark, harder to beat when there is no seasonality; extend the horizon to thirteen weeks to see where the model’s advantage disappears; or redo the exercise with a simpler model, without the second harmonic. Each variant must be announced before reading and reported even if it degrades the verdict.
13 · PROTOCOL
Result interpretation
The model forecasts better than a simple rule, by a wide margin: its error is less than half a unit, while the seasonal naive forecast exceeds one, meaning it does worse out of sample than on its own past. But that comparison is only worth what the benchmark is worth: beating a seasonal naive on a series with a visible trend is not a demanding test. The usable result is therefore modest and precise: on this series and this horizon, this model produced usable forecasts and intervals that are only slightly too narrow.
14 · PROTOCOL
Allowed conclusions
Allowed: using this model to plan four weeks ahead on this series; publishing the scaled error together with the benchmark’s, never one without the other; announcing the 80% intervals while flagging that in practice they cover slightly less; replaying this same validation each quarter to watch for drift; and requiring the same protocol from any forecast supplier. Also allowed: saying that this result holds for this horizon, and redoing the examination if the horizon changes.
15 · PROTOCOL
Forbidden conclusions
Forbidden: presenting 0.420792 as guaranteed accuracy for the weeks ahead; announcing the 80% intervals without mentioning that they covered 0.750000; concluding that a validated model will stay valid after a change of price, channel or competition; comparing this scaled error with that of another series without rechecking that the benchmark is the same; or choosing the number of origins and the horizon after seeing the results. Also forbidden: presenting this synthetic case as an observed measurement.
16 · PROTOCOL
Possible marketing decision
The reasonable decision is to commit the four-week plan to this model, but to size the safety stock on the observed coverage rather than the announced one, since the intervals proved slightly narrow. The validation is then replayed each quarter with the same protocol, unchanged, so that a degradation shows up as a degradation and not as a change of rule. None of these steps automates without someone to read the two flags.
17 · PROTOCOL
When to use or avoid the method
Use this protocol before any forecast goes into production, and at regular intervals afterwards. Avoid it when the history is too short to yield several origins, when past values have been revised after publication, or when the series has just undergone a visible break: in that last case, validation measures a regime that no longer exists. Avoid it too as an answer to the question of an action’s effect, which is a matter for an experiment rather than a forecast.
18 · PROTOCOL
Implementations and final deliverable
The CC0 CSV holds the 156 synthetic weeks. Python and R are the reference implementations; SPSS carries the same computation in a Python block; SAS uses the same design and origins but delegates the fit to its own regression procedure, and checks the two flags rather than every decimal. The final deliverable gathers the data, the variable dictionary, the declared model form, the origins and the horizon, the benchmark, the scaling quantity, the nominal level and the quantile, the two thresholds, the seven numbered steps, the scaled error of the model and of the benchmark, the per-horizon errors, the coverage, the width, the interval score, both flags, the verdict and the software versions.
Synthetic data · CC0
msc-p033-weekly-series.csv ↓Reproducibility protocol
msc-p033-reproducibility-readme.md ↓Python reference · MIT
msc-p033-reference.py ↓R reference · MIT
msc-p033-reference.R ↓SPSS implementation · MIT
msc-p033-secondary.sps ↓SAS implementation · MIT
msc-p033-secondary.sas ↓19 · PROTOCOL
Sources and evidence level
Hyndman and Koehler propose the scaled absolute error as the standard comparison measure, recall that a value above one signals forecasts that are worse on average, and show that many common measures are not generally applicable. Hyndman and Khandakar recall that out-of-sample errors are often too few to conclude, and that two models can give the same point forecasts with different intervals. Makridakis, Spiliotis and Assimakopoulos note the frequent absence of a benchmark and the gap between goodness of fit and out-of-sample accuracy. Petropoulos and co-authors describe the move from a fixed origin to a rolling one, the interval score as the appropriate measure, and the fact that published intervals are too narrow. Founding references: Hyndman and Koehler (2006, International Journal of Forecasting, peer-reviewed) and Hyndman and Khandakar (2008, Journal of Statistical Software, CC BY open access, peer-reviewed). Recent developments: Makridakis, Spiliotis and Assimakopoulos (2018, PLOS ONE, CC BY open access, peer-reviewed) and Petropoulos et al. (2022, arXiv preprint; published in the International Journal of Forecasting, peer-reviewed).
- Hyndman & Koehler (2006) Full text verified, with a short excerpt located in the source.
- Hyndman & Khandakar (2008) Full text verified, with a short excerpt located in the source.
- Makridakis, Spiliotis & Assimakopoulos (2018) Full text verified, with a short excerpt located in the source.
- Petropoulos et al. (2022) Full text verified, with a short excerpt located in the source.
Method connections

