# MSC-P-031 — customer segmentation reproducibility protocol

## Frozen design

- Synthetic data only; no real customer, order or revenue data.
- 900 customers generated with seed `20260911` from four latent groups that differ in recency, frequency, basket, category breadth and digital share.
- Declared features and transformations, fixed before the file is read: log recency, frequency, log basket, category breadth and digital share, each standardized to mean zero and unit standard deviation over the whole file.
- Declared number of segments: four, fixed by the decision context rather than chosen from the data. Choosing the number of segments after looking at the fit is a different exercise and is not what this dossier does.
- Declared construction: Lloyd iterations to at most one hundred passes, starting from the observations located at the ranks 12.5%, 37.5%, 62.5% and 87.5% of a declared composite score, the sum of the standardized features. Nothing is drawn at random, so every implementation reproduces the same digits.
- Declared thresholds, fixed before any result is read: mean silhouette at least 0.25; smallest segment share at least 0.05; adjusted Rand index between two independently rebuilt segmentations at least 0.60; ratio between the highest and lowest segment mean of the outcome at least 1.50.
- The outcome, next-quarter revenue, never enters the construction. It is only read by the usefulness check.
- The verdict is fail-closed: it authorizes acting on the segments only when the four checks pass.

## Expected output

`customers=900`, `segments=4`, `segment_sizes=283,246,184,187`, `segment_shares=0.314444,0.273333,0.204444,0.207778`, `smallest_share=0.204444`, `balance_flag=PASS`, `mean_silhouette=0.407973`, `separation_flag=PASS`, `stability_adjusted_rand=0.984661`, `stability_flag=PASS`, `segment_outcome_means=192.210318,106.899593,69.260000,644.625668`, `outcome_ratio_high_low=9.307330`, `outcome_variance_explained=0.692635`, `usefulness_flag=PASS`, `verdict=SEGMENTATION_READABLE_FOR_ACTION`, plus one `centroid_k` line per segment.

## Run

```text
python msc-p031-reference.py msc-p031-segmentation-panel.csv
Rscript msc-p031-reference.R msc-p031-segmentation-panel.csv
```

The SPSS syntax carries its computation in a `BEGIN PROGRAM Python3` block; that block was executed outside SPSS on the shipped CSV and printed the same values as the Python reference. The SAS program is different by design: PROC FASTCLUS receives the same four declared seeds, but its assignment order, tie-breaking and convergence rule are its own, and it checks balance and usefulness rather than reproducing every digit. The SPSS, SAS and R runtimes are not available in the local verification environment, so those three files are reviewed statically and must be run independently before relying on their printed output.

## Variable dictionary

| Column | Type | Unit | Definition |
|---|---|---|---|
| `customer_id` | text, `S0001`…`S0900` | — | Customer identifier, ordered without gaps. |
| `recency_days` | number > 0 | days | Days since the last purchase; entered as its logarithm. |
| `frequency_12m` | number > 0 | orders | Orders over the last twelve months. |
| `avg_basket_eur` | number > 0 | currency units | Average basket over the same period; entered as its logarithm. |
| `category_breadth` | number > 0 | categories | Average number of distinct categories bought per period. |
| `digital_share` | number in (0, 1) | share | Share of orders placed online. |
| `next_quarter_revenue_eur` | number ≥ 0 | currency units | Revenue observed in the following quarter. Never used to build the segments. |

Derived quantities printed by the reference scripts:

| Output | Definition |
|---|---|
| `segment_sizes`, `segment_shares`, `smallest_share` | Number and share of customers per segment, and the smallest share. |
| `balance_flag` | `PASS` when the smallest share reaches 0.05, so that no segment is too small to address. |
| `mean_silhouette` | Mean silhouette over a declared systematic sample of every third customer, each compared with all customers of every segment. |
| `separation_flag` | `PASS` when that mean reaches 0.25. |
| `stability_adjusted_rand` | Adjusted Rand index between the labels obtained from two segmentations rebuilt independently on the odd-numbered and even-numbered customers, both applied to all customers. |
| `stability_flag` | `PASS` when that index reaches 0.60. |
| `segment_outcome_means`, `outcome_ratio_high_low`, `outcome_variance_explained` | Mean next-quarter revenue per segment, the ratio between the highest and the lowest, and the share of the outcome variance the segmentation accounts for. |
| `usefulness_flag` | `PASS` when the ratio reaches 1.50. |
| `centroid_k` | Standardized profile of segment k, in the declared feature order. |
| `verdict` | `SEGMENTATION_READABLE_FOR_ACTION` only when the four flags pass; otherwise `DIAGNOSTIC_BLOCKS_SEGMENTATION_READING`. |

## Numbered analysis steps

1. Load the CSV; refuse any file whose header, cell count, customer order, positivity or share bounds differ from the frozen schema.
2. Apply the declared transformations and standardize the five features over the whole file.
3. Place the four initial centroids at the declared ranks of the composite score and run Lloyd iterations until the assignment stops changing or one hundred passes are reached; fail if a segment empties.
4. Report the sizes and shares and apply the balance threshold.
5. Compute the mean silhouette over the declared systematic sample and apply the separation threshold.
6. Rebuild the segmentation independently on the odd-numbered and on the even-numbered customers, label all customers under each solution, compute the adjusted Rand index between the two labelings and apply the stability threshold.
7. Compute the mean outcome per segment, the ratio between the highest and lowest, and the share of variance explained, apply the usefulness threshold, then apply the fail-closed verdict rule and print every quantity of the expected output; a third party compares the printed lines with the expected output above, digit for digit.

## Scientific boundary

A segmentation is a description of the data under a declared distance, not a discovery of natural groups. The same customers, described by other features or another distance, would give other segments; the four checks say whether these segments hold together and separate the outcome, never whether they are the true ones. The silhouette here is computed on a systematic sample of customers but against all customers, which is not identical to a silhouette computed inside the sample alone; a standard library that restricts distances to the sampled points returns a slightly different value, 0.406210 instead of 0.407973 on this file. The stability check compares two rebuilds on halves of the same file; it does not test stability over time, which is a different question. The usefulness check reads an outcome that the construction never saw, but a difference in outcome across segments is an association, not the effect of treating a segment differently. Nothing here identifies what would happen if a segment received a specific offer, and nothing transports to another period or market.

Dataset license: CC0-1.0. Code license: MIT.
