Research & Evidence

Search for a method

Search titles, questions, territories and MSC identifiers.

41 results
  1. MSC-P-001How do you turn a marketing claim into a testable question?Decision Science↗
  2. MSC-P-002Correlation or causality: what can an analysis actually support?Marketing Measurement↗
  3. MSC-P-003How should uncertainty in a marketing result be expressed?Decision Science↗
  4. MSC-P-004Statistical significance or effect size: which result should be interpreted?Decision Science↗
  5. MSC-P-005How do you measure a marketing construct that is not directly observable?Market Research↗
  6. MSC-P-006How do you design and validate a measurement scale?Market Research↗
  7. MSC-P-007Alpha or omega: how should scale reliability be assessed?Market Research↗
  8. MSC-P-009PCA, EFA or CFA: which method should you choose?Market Research↗
  9. MSC-P-010When should you run a marketing experiment?Marketing Measurement↗
  10. MSC-P-011How do you design an A/B test that actually estimates an effect?Marketing Measurement↗
  11. MSC-P-012How many observations does an experiment need?Decision Science↗
  12. MSC-P-013How do you measure campaign incrementality with a control group?Marketing Measurement↗
  13. MSC-P-017How do you detect selection, contamination and attrition in an experiment?Marketing Measurement↗
  14. MSC-P-018Predictive or causal regression: what are you trying to estimate?Marketing Models↗
  15. MSC-P-019How do you diagnose a marketing regression before interpreting it?Marketing Models↗
  16. MSC-P-022How do you estimate price elasticity and its uncertainty?Pricing Science↗
  17. MSC-P-026Logit vs Probit: how do you choose for purchase probability?Customer Science↗
  18. MSC-P-029Which customers have the highest probability of churn?Customer Science↗
  19. MSC-P-027TAM, UTAUT or UTAUT2: which framework should be used to study technology acceptance?Market Research↗
  20. MSC-H-001Measurement and causality: how can a marketing effect be established?Marketing Measurement↗
  21. MSC-H-002Marketing response models: shape, delay and saturationMarketing Models↗
  22. MSC-H-003Pricing science: connecting price, demand and contributionPricing Science↗
  23. MSC-H-004Customer and choice science: behavior, value and heterogeneityCustomer Science↗
  24. MSC-H-005Measurement science: building valid indicatorsMarket Research↗
  25. MSC-H-006Statistical decision methods: choose, quantify, validateDecision Science↗
  26. MSC-P-008How do you validate a marketing measurement scale?Market Research↗
  27. MSC-P-014How do you design a marketing geo experiment?Marketing Measurement↗
  28. MSC-P-015How do you estimate an effect with difference-in-differences?Marketing Measurement↗
  29. MSC-P-020How do you address price endogeneity?Pricing Science↗
  30. MSC-P-021Fixed or random effects: which panel model should you choose?Marketing Models↗
  31. MSC-P-023How do you estimate a demand function?Pricing Science↗
  32. MSC-P-024How do you simulate a price-volume-margin scenario?Pricing Science↗
  33. MSC-P-028How do you estimate CLV with BG/NBD and Gamma-Gamma?Customer Science↗
  34. MSC-P-030How do you analyze retention with a survival model?Customer Science↗
  35. MSC-P-043What is a subscriber worth with only six months of retention data?Customer Science↗
  36. MSC-P-031How do you build a useful customer segmentation?Customer Science↗
  37. MSC-P-032How do you test segmentation stability?Customer Science↗
  38. MSC-P-033How do you validate a marketing forecast?Decision Science↗
  39. MSC-P-034How do you build a Monte Carlo simulation for a marketing decision?Decision Science↗
  40. MSC-P-035How do you model saturation and adstock?Marketing Models↗
  41. MSC-P-039Which statistical test should you choose?Decision Science↗
← All methods
METHOD DOSSIERMSC-P-033Regression and econometricsVerified scientific dossier

How do you validate a marketing forecast?

A forecast is validated on periods not used for fitting, with a realistic forecast origin, a naive benchmark and a metric aligned with error cost.

Scientific editorial team: Marketing Science Center

Direct answer

Compare models for a declared horizon and error cost.

A forecast is validated on periods not used for fitting, with a realistic forecast origin, a naive benchmark and a metric aligned with error cost.

Shmueli, 2010Hyndman & Koehler, 2006

01 · PROTOCOL

Operational summary

A weekly forecasting model is refitted at thirteen successive origins, then judged on 52 weeks it has never seen. Its scaled absolute error is 0.420792, against 1.039025 for a seasonal naive forecast: it clearly beats the benchmark. Its 80% intervals contain the outcome 0.750000 of the time, inside the declared tolerance of 0.10 but on the narrow side. Both checks pass. Verdict: FORECAST_READABLE_FOR_PLANNING.

02 · PROTOCOL

Concrete marketing situation

A team must commit stock and media budgets four weeks ahead. It has a model that forecasts weekly revenue and that, on the history, fits the data remarkably well. The question to ask before using it is simple: when has this model ever faced weeks it did not know, and what happened? Without that answer, the observed goodness of fit says nothing about the intended use.

03 · PROTOCOL

Scientific question

For this weekly series, this declared model, these thirteen declared origins and this four-week horizon, is the model’s error on weeks not used to fit it lower than that of a seasonal naive forecast brought to the same scale? And do its 80% intervals contain the outcome as often as they claim? The two questions are distinct and each gets its own answer: an accurate point forecast can come with a false uncertainty.

04 · PROTOCOL

Why the simple approach can fail

Judging a model on the weeks used to fit it always flatters: the more parameters one adds, the better the result looks, without any forecasting ability having been demonstrated. Splitting the history once into training and test is not enough either, since the score then depends on the accident of a single boundary. And reporting an error in currency units with no benchmark leaves nobody able to judge: a mean error of 55 units is excellent or catastrophic depending on the series.

05 · PROTOCOL

Method intuition

The method reproduces real use. One steps back to a past date, uses only what was known that day, forecasts the following four weeks, then moves four weeks forward and starts again. Thirteen times. Each error is then divided by a reference quantity computed on the past available at that date alone: the error a seasonal naive forecast would have made. A value below one means the model did better than that simple rule. In parallel, one counts how often the announced interval actually contained the outcome.

06 · PROTOCOL

Required data

Frozen before any reading: the gapless week numbering, the weekly revenue, the model form (constant, linear trend and two annual harmonics), the seasonal period of 52 weeks, the thirteen origins from 104 to 152 in steps of 4, the horizon of 4 weeks, the seasonal naive forecast as benchmark, the scaling quantity, the nominal level of 0.80, the quantile 1.281552 and the two thresholds. The sealed file holds 156 weeks. The case is synthetic and declared as such.

07 · PROTOCOL

Formal model and symbols

At each origin, the model fits by least squares a constant, a linear trend and two annual harmonics — six parameters — on the weeks before the origin only. The scaled absolute error divides each out-of-sample error by the mean of the 52-week absolute differences observed in that same window. The interval has a half-width of 1.281552 times the residual standard deviation of the fit, that quantile being the 0.90 quantile of the standard normal law. The interval score adds to the width a penalty proportional to any miss, divided by the accepted risk.

08 · PROTOCOL

Declared calculation

The declared computation runs seven steps: validate the schema and refuse any non-conforming file or one too short for the validation design; at each origin, build the design matrix on the earlier weeks only; solve the normal equations by Gaussian elimination with partial pivoting, refuse a singular window and compute the residual standard deviation with the six parameters deducted from the degrees of freedom; compute the seasonal scale on the same window and refuse a zero scale; for each of the four horizons, record the absolute error, the scaled error, the naive one, whether the interval contains the outcome, the width and the score; average over the 52 forecasts and then separately by horizon; apply the two declared thresholds, then the verdict rule.

09 · PROTOCOL

End-to-end numeric example

Synthetic illustration — teaching values, not observed

On the sealed file: 156 weeks, 13 origins from week 104 to week 152, horizon 4, hence 52 out-of-sample forecasts. Model scaled absolute error 0.420792, against 1.039025 for the seasonal naive forecast; mean absolute error 54.912955 currency units; accuracy check passed. By horizon: 0.418133, 0.520418, 0.279945 and 0.464672. Nominal level 0.80, observed coverage 0.750000, mean width 168.860643 and mean interval score 243.378776; calibration check passed. Verdict: FORECAST_READABLE_FOR_PLANNING.

10 · PROTOCOL

Validity assumptions

Validation assumes that the mechanism which produced past weeks is still the one that will produce future weeks, that the seasonal period really is 52 weeks, and that historical values have not been revised after the fact. It also assumes, for the intervals, that the residuals behave like normal noise with constant variance. Three things are not testable here: the absence of a future regime change, the effect of a channel or a price that never varied in the history, and the quality of the data at the moment the forecast will actually be produced.

11 · PROTOCOL

Diagnostics and uncertainty

Two caveats deserve to be read before the verdict. The observed coverage is 0.750000 for an announced level of 0.80: the gap stays inside the declared tolerance, but the intervals are on the narrow side, which is exactly what one expects from intervals built on residual spread alone, ignoring both parameter uncertainty and its growth with the horizon. And the per-horizon errors — 0.418133, 0.520418, 0.279945, 0.464672 — do not increase steadily with the horizon: with thirteen forecasts per horizon that pattern is noise, and no threshold is attached to them.

12 · PROTOCOL

Robustness and alternatives

Alternatives declared before results: replace the normal intervals by intervals built on the empirical quantiles of the out-of-sample errors, one set per horizon; add a simple random walk to the benchmark, harder to beat when there is no seasonality; extend the horizon to thirteen weeks to see where the model’s advantage disappears; or redo the exercise with a simpler model, without the second harmonic. Each variant must be announced before reading and reported even if it degrades the verdict.

13 · PROTOCOL

Result interpretation

The model forecasts better than a simple rule, by a wide margin: its error is less than half a unit, while the seasonal naive forecast exceeds one, meaning it does worse out of sample than on its own past. But that comparison is only worth what the benchmark is worth: beating a seasonal naive on a series with a visible trend is not a demanding test. The usable result is therefore modest and precise: on this series and this horizon, this model produced usable forecasts and intervals that are only slightly too narrow.

14 · PROTOCOL

Allowed conclusions

Allowed: using this model to plan four weeks ahead on this series; publishing the scaled error together with the benchmark’s, never one without the other; announcing the 80% intervals while flagging that in practice they cover slightly less; replaying this same validation each quarter to watch for drift; and requiring the same protocol from any forecast supplier. Also allowed: saying that this result holds for this horizon, and redoing the examination if the horizon changes.

15 · PROTOCOL

Forbidden conclusions

Forbidden: presenting 0.420792 as guaranteed accuracy for the weeks ahead; announcing the 80% intervals without mentioning that they covered 0.750000; concluding that a validated model will stay valid after a change of price, channel or competition; comparing this scaled error with that of another series without rechecking that the benchmark is the same; or choosing the number of origins and the horizon after seeing the results. Also forbidden: presenting this synthetic case as an observed measurement.

16 · PROTOCOL

Possible marketing decision

The reasonable decision is to commit the four-week plan to this model, but to size the safety stock on the observed coverage rather than the announced one, since the intervals proved slightly narrow. The validation is then replayed each quarter with the same protocol, unchanged, so that a degradation shows up as a degradation and not as a change of rule. None of these steps automates without someone to read the two flags.

17 · PROTOCOL

When to use or avoid the method

Use this protocol before any forecast goes into production, and at regular intervals afterwards. Avoid it when the history is too short to yield several origins, when past values have been revised after publication, or when the series has just undergone a visible break: in that last case, validation measures a regime that no longer exists. Avoid it too as an answer to the question of an action’s effect, which is a matter for an experiment rather than a forecast.

18 · PROTOCOL

Implementations and final deliverable

The CC0 CSV holds the 156 synthetic weeks. Python and R are the reference implementations; SPSS carries the same computation in a Python block; SAS uses the same design and origins but delegates the fit to its own regression procedure, and checks the two flags rather than every decimal. The final deliverable gathers the data, the variable dictionary, the declared model form, the origins and the horizon, the benchmark, the scaling quantity, the nominal level and the quantile, the two thresholds, the seven numbered steps, the scaled error of the model and of the benchmark, the per-horizon errors, the coverage, the width, the interval score, both flags, the verdict and the software versions.

19 · PROTOCOL

Sources and evidence level

Hyndman and Koehler propose the scaled absolute error as the standard comparison measure, recall that a value above one signals forecasts that are worse on average, and show that many common measures are not generally applicable. Hyndman and Khandakar recall that out-of-sample errors are often too few to conclude, and that two models can give the same point forecasts with different intervals. Makridakis, Spiliotis and Assimakopoulos note the frequent absence of a benchmark and the gap between goodness of fit and out-of-sample accuracy. Petropoulos and co-authors describe the move from a fixed origin to a rolling one, the interval score as the appropriate measure, and the fact that published intervals are too narrow. Founding references: Hyndman and Koehler (2006, International Journal of Forecasting, peer-reviewed) and Hyndman and Khandakar (2008, Journal of Statistical Software, CC BY open access, peer-reviewed). Recent developments: Makridakis, Spiliotis and Assimakopoulos (2018, PLOS ONE, CC BY open access, peer-reviewed) and Petropoulos et al. (2022, arXiv preprint; published in the International Journal of Forecasting, peer-reviewed).

  1. Hyndman & Koehler (2006) Full text verified, with a short excerpt located in the source.
  2. Hyndman & Khandakar (2008) Full text verified, with a short excerpt located in the source.
  3. Makridakis, Spiliotis & Assimakopoulos (2018) Full text verified, with a short excerpt located in the source.
  4. Petropoulos et al. (2022) Full text verified, with a short excerpt located in the source.

Method connections

Parent territoryStatistical decision methods: choose, quantify, validateRequiresHow should uncertainty in a marketing result be expressed?Compare withFixed or random effects: which panel model should you choose?

Read next

MSC-H-002Marketing response models: shape, delay and saturation→MSC-H-006Statistical decision methods: choose, quantify, validate→MSC-P-003How should uncertainty in a marketing result be expressed?→MSC-P-021Fixed or random effects: which panel model should you choose?→