# MSC-P-033 — marketing forecast validation reproducibility protocol

## Frozen design

- Synthetic data only; no real revenue or campaign data.
- 156 weekly observations generated with seed `20260913` from a level, a linear trend, an annual seasonality carried by two Fourier harmonics at period 52, and Gaussian noise.
- Declared model, fixed before the file is read: least squares on an intercept, a linear trend and two annual harmonics — six parameters, refitted from scratch at every forecast origin on the weeks available up to that origin. No week beyond an origin ever enters its own fit.
- Declared benchmark: the seasonal naive forecast, that is the value observed 52 weeks earlier.
- Declared validation design: forecast origins at weeks 104, 108, … 152, thirteen in all, each followed over a horizon of 4 weeks, for 52 out-of-sample forecasts. This is a rolling origin, not a single train/test split.
- Declared scale for the mean absolute scaled error: at each origin, the in-sample mean absolute seasonal naive error over the weeks available up to that origin. Scaling by an in-sample quantity is what makes the measure comparable across series and keeps the benchmark from being judged on its own errors.
- Declared intervals: nominal level 0.80, half-width `1.281552 × s`, where `1.281552` is the 0.90 quantile of the standard normal law and `s` is the residual standard deviation of the fit at that origin, with the six estimated parameters deducted from the degrees of freedom. This is the interval a practitioner builds by default; the dossier tests it rather than assuming it.
- Declared thresholds, fixed before any result is read: the model's mean absolute scaled error must be strictly below 1.00, that is it must beat the seasonal naive benchmark on the same scale; and the empirical coverage of the intervals must lie within 0.10 of their nominal 0.80, that is between 0.70 and 0.90.
- Errors are also reported by horizon and as a mean interval score, without thresholds, because thirteen forecasts per horizon are too few to carry one.
- The verdict is fail-closed: it authorizes planning on the model only when both checks pass.

## Expected output

`origins=13`, `first_origin=104`, `last_origin=152`, `horizon=4`, `forecasts=52`,
`model_mase=0.420792`, `benchmark_mase=1.039025`, `model_mae=54.912955`, `accuracy_flag=PASS`,
`mase_h1=0.418133`, `mase_h2=0.520418`, `mase_h3=0.279945`, `mase_h4=0.464672`,
`nominal_coverage=0.80`, `empirical_coverage=0.750000`, `mean_interval_width=168.860643`, `mean_interval_score=243.378776`, `calibration_flag=PASS`,
`verdict=FORECAST_READABLE_FOR_PLANNING`.

## Run

```text
python msc-p033-reference.py msc-p033-weekly-series.csv
Rscript msc-p033-reference.R msc-p033-weekly-series.csv
```

The SPSS syntax carries its computation in a `BEGIN PROGRAM Python3` block; that block was executed outside SPSS on the shipped CSV and printed the same values as the Python reference. The SAS program uses the same declared design and origins but delegates the fit to PROC REG, whose solver and residual conventions are its own; it checks the two flags and the verdict rather than every digit. The SPSS, SAS and R runtimes are not available in the local verification environment, so those three files are reviewed statically and must be run independently before relying on their printed output.

## Variable dictionary

| Column | Type | Unit | Definition |
|---|---|---|---|
| `week` | integer, 1…156 | week | Week number, ordered without gaps. |
| `revenue_eur` | number > 0 | currency units | Revenue observed that week. |

Derived quantities printed by the reference scripts:

| Output | Definition |
|---|---|
| `origins`, `first_origin`, `last_origin`, `horizon`, `forecasts` | The declared rolling-origin design and the number of out-of-sample forecasts it produces. |
| `model_mase` | Mean absolute scaled error of the model: each out-of-sample absolute error divided by the in-sample seasonal naive scale of its own origin, averaged over the 52 forecasts. Below 1 means the model beats the seasonal naive benchmark. |
| `benchmark_mase` | The same quantity for the seasonal naive forecast, reported so that the comparison is visible rather than asserted. |
| `model_mae` | Mean absolute error in currency units, which is readable but not comparable across series. |
| `accuracy_flag` | `PASS` when `model_mase < 1.00`. |
| `mase_h1`…`mase_h4` | The same scaled error restricted to each horizon, reported without a threshold. |
| `nominal_coverage`, `empirical_coverage` | The level the intervals claim, and the share of the 52 outcomes they actually contain. |
| `mean_interval_width` | Mean width of the intervals, in currency units. An interval can reach its nominal coverage simply by being wide. |
| `mean_interval_score` | Mean Winkler interval score: the width plus a penalty proportional to any miss. It rewards intervals that are both narrow and honest. |
| `calibration_flag` | `PASS` when the empirical coverage lies within 0.10 of the nominal level. |
| `verdict` | `FORECAST_READABLE_FOR_PLANNING` only when both flags pass; otherwise `DIAGNOSTIC_BLOCKS_FORECAST_READING`. |

## Numbered analysis steps

1. Load the CSV; refuse any file whose header, cell count, week numbering or positivity differ from the frozen schema, or that is shorter than the declared validation design.
2. For each declared origin, build the declared design matrix on weeks 1 to that origin only.
3. Solve the normal equations by Gaussian elimination with partial pivoting; refuse a singular window; compute the residual standard deviation with the six parameters deducted from the degrees of freedom.
4. Compute the declared seasonal scale on the same window, and refuse a zero scale.
5. For each of the four horizons, predict the week, record the absolute error, the scaled error, the seasonal naive scaled error, whether the interval contains the outcome, the interval width and the interval score.
6. Average across the 52 forecasts, and separately by horizon.
7. Apply the two declared thresholds to obtain both flags, apply the fail-closed verdict rule, and print every quantity of the expected output; a third party compares the printed lines with the expected output above, digit for digit.

## Scientific boundary

A validation says how a model behaved on weeks it had not seen, in a period that has already happened. It does not promise the same behaviour after a change of regime, a new channel, a price change or a competitor's move, and no out-of-sample score can be turned into such a promise. Three limits are worth stating plainly. The point forecast is clearly better than the seasonal naive benchmark here, but that comparison is only as meaningful as the benchmark: beating a seasonal naive on a series with a visible trend is not a demanding test. The intervals pass their check while covering 0.750000 rather than 0.80 — inside the declared tolerance, but on the narrow side, which is what one expects from intervals built on residual spread alone, ignoring both parameter uncertainty and the growth of uncertainty with the horizon. And the per-horizon errors do not increase monotonically with the horizon on this file; with thirteen forecasts per horizon that pattern is noise, which is why no threshold is attached to them. Finally, the case is synthetic: the model form used to forecast is the form that generated the data, so this dossier demonstrates that the validation procedure behaves as declared, not that a real marketing series is this predictable.

Dataset license: CC0-1.0. Code license: MIT.
