# MSC-P-032 — segmentation stability reproducibility protocol

## Frozen design

- Synthetic data only; no real customer or order data.
- 360 customers generated with seed `20260912` from **three** latent groups of unequal size (150, 120, 90) that differ in recency, frequency and basket. The question put to the dossier is whether a **four**-segment solution proposed by a business survives resampling; the file was built with three groups on purpose, and whatever the checks returned is what is published.
- Declared features and transformations, fixed before the file is read: log recency, frequency and log basket, each standardized to mean zero and unit standard deviation over the whole file.
- Declared clustering: Lloyd iterations to at most one hundred passes, starting from the observations at the ranks (j+0.5)/k of a declared composite score, the sum of the standardized features. Nothing is drawn from a library generator.
- Declared resampling: 40 bootstrap resamples of the same size drawn **with replacement** by the linear congruential recursion `state ← (1103515245 · state + 12345) mod 2^31`, seeded at `20260912`, with `index = state mod n`. The recursion is arithmetic, so Python, R, SPSS and SAS draw the identical resamples without any seed convention. One caveat matters for implementers: written directly, the product `1103515245 · state` reaches 2.4e18, well above 2^53, the largest integer a double holds exactly. Languages whose numbers are doubles — R and SAS among them — must therefore split the multiplier as `16838 · 65536 + 20077`, which keeps every intermediate product exact. The R and SAS files shipped here do so; they were corrected no later than 2026-09-06 (the record first read 2026-09-15, a day that had not yet come), after the defect was found while building MSC-P-039.
- Declared comparison: each resample is clustered, then **all** original customers are relabelled under the resample's centroids, so both partitions cover the same objects and the Jaccard index is defined without matching sets of different sizes. Each original segment is scored by its best Jaccard against the rebuilt segments, and the score is averaged over the 40 resamples.
- Declared structureless reference: each feature is permuted independently by the same recursion, destroying any joint structure while preserving each marginal distribution. The identical procedure is run on it, so a recovery level is read against what resampling noise alone already produces.
- Declared thresholds, fixed before any result is read: a solution is stable when the weakest segment recovery reaches 0.75 **and** exceeds the structureless reference for the same number of segments by at least 0.10.
- Candidate numbers of segments reported: 2, 3 and 4. The verdict concerns only the proposed number, 4; the other two are reported so that the reader sees what stability does and does not settle.
- The verdict is fail-closed: it authorizes acting on the proposed segmentation only when both thresholds are met.

## Expected output

`customers=360`, `resamples=40`, `proposed_segments=4`,
`recovery_k2=0.977481,0.976322`, `weakest_k2=0.976322`, `null_weakest_k2=0.647971`, `margin_k2=0.328351`, `flag_k2=STABLE`,
`recovery_k3=0.990139,0.986949,0.997222`, `weakest_k3=0.986949`, `null_weakest_k3=0.606088`, `margin_k3=0.380860`, `flag_k3=STABLE`,
`recovery_k4=0.848874,0.664749,0.827341,0.979621`, `weakest_k4=0.664749`, `null_weakest_k4=0.829190`, `margin_k4=-0.164441`, `flag_k4=NOT_STABLE`,
`verdict=PROPOSED_SEGMENTATION_NOT_STABLE`.

## Run

```text
python msc-p032-reference.py msc-p032-stability-panel.csv
Rscript msc-p032-reference.R msc-p032-stability-panel.csv
```

The SPSS syntax carries its computation in a `BEGIN PROGRAM Python3` block; that block was executed outside SPSS on the shipped CSV and printed the same values as the Python reference. The SAS program reproduces the declared recursion exactly, so it draws the same resamples, but it delegates the clustering to PROC FASTCLUS, whose assignment order and convergence rule are its own; it checks the flags and the verdict rather than every digit. The SPSS, SAS and R runtimes are not available in the local verification environment, so those three files are reviewed statically and must be run independently before relying on their printed output.

## Variable dictionary

| Column | Type | Unit | Definition |
|---|---|---|---|
| `customer_id` | text, `T0001`…`T0360` | — | Customer identifier, ordered without gaps. |
| `recency_days` | number > 0 | days | Days since the last purchase; entered as its logarithm. |
| `frequency_12m` | number > 0 | orders | Orders over the last twelve months. |
| `avg_basket_eur` | number > 0 | currency units | Average basket over the same period; entered as its logarithm. |

Derived quantities printed by the reference scripts:

| Output | Definition |
|---|---|
| `recovery_kK` | Mean Jaccard recovery of each segment of the K-segment solution over the 40 declared resamples, in the order of the original segments. |
| `weakest_kK` | The smallest of those recoveries: a segmentation is only as trustworthy as its least reproducible segment. |
| `null_weakest_kK` | The same quantity computed on the structureless reference, that is, the recovery that resampling noise alone produces at K segments. |
| `margin_kK` | `weakest_kK − null_weakest_kK`. A negative margin means the real file reproduces its segments **less** well than a file with no joint structure. |
| `flag_kK` | `STABLE` when `weakest_kK ≥ 0.75` and `margin_kK ≥ 0.10`; otherwise `NOT_STABLE`. |
| `verdict` | `PROPOSED_SEGMENTATION_STABLE` only when the flag of the proposed number of segments is `STABLE`; otherwise `PROPOSED_SEGMENTATION_NOT_STABLE`. |

## Numbered analysis steps

1. Load the CSV; refuse any file whose header, cell count, customer order or positivity differ from the frozen schema.
2. Apply the declared transformations and standardize the three features over the whole file.
3. Build the declared structureless reference by permuting each standardized feature independently with the declared recursion.
4. For each candidate number of segments, build the declared clustering on the whole file and record the composition of each segment.
5. Draw the 40 declared resamples with replacement, cluster each resample, relabel every original customer under the resample's centroids, and score each original segment by its best Jaccard against the rebuilt segments; fail if a resample cannot produce the requested number of segments.
6. Average each segment's score over the resamples, repeat steps 4 and 5 unchanged on the structureless reference, and compute the weakest recovery, the reference weakest recovery and their margin.
7. Apply the two declared thresholds to obtain each flag, apply the fail-closed verdict rule to the proposed number of segments, and print every quantity of the expected output; a third party compares the printed lines with the expected output above, digit for digit.

## Scientific boundary

This protocol measures whether a segmentation reappears when the same customers are resampled. It does not measure whether the segments are true, useful, or distinct in their response to an action. Three limits are structural. First, stability does not select the number of segments: on this file the two-segment solution is also stable, because merging genuine groups yields a solution that resampling reproduces easily. Second, a high recovery is not by itself evidence of structure — at four segments the structureless reference recovers its own segments better (0.829190) than the real file recovers its own (0.664749), which is exactly why the reference is computed. Third, resampling the same file says nothing about a different period, a different market, or a change in the variables themselves; and nothing here identifies what would happen if a segment received a specific offer.

Dataset license: CC0-1.0. Code license: MIT.
