How do you build a useful customer segmentation?
A useful segmentation connects admissible variables, distance measure, algorithm, stability and decision use. Groups are not natural essences but a conditional representation.
Scientific editorial team: Marketing Science Center
Direct answer
Build an actionable descriptive partition and document its uncertainty.
A useful segmentation connects admissible variables, distance measure, algorithm, stability and decision use. Groups are not natural essences but a conditional representation.
01 · PROTOCOL
Operational summary
Nine hundred customers are split into four segments by a fully declared construction, with no random draw. In the sealed synthetic case the sizes are 283, 246, 184 and 187, the mean silhouette 0.407973 and the adjusted Rand index between two independent rebuilds 0.984661. Next-quarter revenue, never used to build the segments, runs from 69.260000 to 644.625668 across segments, a ratio of 9.307330. The four checks pass. Verdict: SEGMENTATION_READABLE_FOR_ACTION.
02 · PROTOCOL
Concrete marketing situation
A customer relationship team wants to stop treating its base as a single block. For each customer it has recency, frequency, average basket, number of categories bought and share of online orders. It wants four groups, because four is the number of relationship programmes it can actually run. The question is not to find the true groups, which do not exist as such, but to obtain a split that holds together and helps decide.
03 · PROTOCOL
Scientific question
For these nine hundred customers, these five declared variables and this declared distance, is the four-segment split separated, balanced, reproducible on other observations, and does it distinguish an outcome the construction never saw? The question is not whether natural groups exist in the customer base, but whether a declared split holds together and is useful for a decision.
04 · PROTOCOL
Why the simple approach can fail
Running a partitioning algorithm and commenting on the groups always produces groups: the method manufactures as many as it is asked for, even in a cloud without structure. Changing the variables, the scale or the distance changes the result, and a different initial draw can move the boundaries. Without checks, one describes an artefact of the algorithm and calls it a customer segmentation. The phrase “true segment” has no operational definition in the literature either.
05 · PROTOCOL
Method intuition
The method consists in declaring everything before looking: the five variables, their transformation, their scaling, the distance, the number of segments, the starting point of the centres and the four thresholds. The start is not drawn at random: it is the observations located at four fixed ranks of a declared composite score. The split is then submitted to four questions: are the segments separated, are they all addressable, do they reappear when rebuilt on other customers, and do they distinguish an outcome outside the construction?
06 · PROTOCOL
Required data
Frozen before any reading: the customer identifier ordered without gaps, recency, frequency, average basket, number of categories, online share, next-quarter revenue, the declared transformations, the standardization, the number of segments fixed at four by the decision context, the declared start and the four thresholds. Revenue never enters the construction. The sealed file holds nine hundred customers. The case is synthetic and declared as such.
07 · PROTOCOL
Formal model and symbols
Each customer is a vector of five standardized coordinates: log recency, frequency, log basket, number of categories, online share. The distance is squared Euclidean. A segment is a set of customers closer to their centre than to any other; the centre is the mean of its members. A customer’s silhouette compares the mean distance to the members of its own segment with the mean distance to the nearest segment. The adjusted Rand index compares two splits while correcting for agreement due to chance.
08 · PROTOCOL
Declared calculation
The declared computation runs seven steps: validate the schema and refuse any non-conforming file; transform and standardize the five variables over the whole file; place the four initial centres at the declared ranks of the composite score and iterate until assignments stop changing or one hundred passes, failing if a segment empties; report sizes and shares and apply the balance threshold; compute the mean silhouette on the declared systematic sample; rebuild separately on even- and odd-numbered customers, relabel the whole file under each solution and compute the adjusted Rand index; compute the outcome means per segment, their ratio and the variance explained, then apply the verdict rule.
09 · PROTOCOL
End-to-end numeric example
Synthetic illustration — teaching values, not observed
On the sealed file: 900 customers, 4 segments. Sizes 283, 246, 184, 187; shares 0.314444, 0.273333, 0.204444 and 0.207778; smallest share 0.204444, check passed. Mean silhouette 0.407973, check passed. Adjusted Rand index between the two rebuilds 0.984661, check passed. Mean revenue per segment: 192.210318; 106.899593; 69.260000; 644.625668. Ratio of highest to lowest 9.307330 and variance explained 0.692635, check passed. Verdict: SEGMENTATION_READABLE_FOR_ACTION.
10 · PROTOCOL
Validity assumptions
The construction assumes that the five variables describe what the team wants to distinguish, that standardization gives them an acceptable weight, that Euclidean distance is meaningful on these coordinates, and that four segments match a real capacity to act. It does not assume that natural groups exist. Three things are not testable here: the relevance of the chosen variables, the stability of the split over time, and whether a segment reacts differently to an action, which is a matter for an experiment.
11 · PROTOCOL
Diagnostics and uncertainty
The silhouette measures a relative separation under the declared distance, not a geometric truth: it would differ with other variables. It is computed here on a systematic sample of customers, but against all customers; a standard library that restricts distances to the sampled points alone returns 0.406210 instead of 0.407973 on this file, and the page says so rather than implying a single definition. The adjusted Rand index reaches 0.984661 because the synthetic structure is sharp; on real data, a value above 0.60 is already reassuring. None of these numbers is a p-value.
12 · PROTOCOL
Robustness and alternatives
Alternatives declared before results: rebuild the split with three or five segments and compare the four checks; replace standardization by an explicit weighting of the variables and publish both profiles; drop the online share, which may reflect an acquisition channel rather than a behaviour; or replace the partition by a method that allows partial membership. Each variant must be announced before reading and reported even if it changes the segments.
13 · PROTOCOL
Result interpretation
The four checks pass, so the split may serve for deciding. The usable result is the profile of the segments and the value gap between them: one segment concentrates a mean revenue of 644.625668 over the next quarter, another 69.260000, and the split accounts for 69.2635 percent of the revenue variance. That structure justifies differentiated programmes. It does not say that a programme will make one segment react more than another: this is an association, measured on an outcome the construction never saw, not an effect.
14 · PROTOCOL
Allowed conclusions
Allowed: publish the segment profiles, their sizes, their separation, their reproducibility and their value gap; build four distinct relationship programmes; track each segment’s share over time; conclude that the base is not homogeneous and that four groups hold under the declared criteria. Also allowed: say that these segments depend on the chosen variables, and republish the split if those variables change.
15 · PROTOCOL
Forbidden conclusions
Forbidden: presenting these four segments as the true groups of the customer base; asserting that a customer belongs to a segment for good, when their position moves as soon as their purchases change; attributing the revenue gap to segment membership as if it were an effect; choosing the number of segments after seeing the profiles and then publishing the checks as if it had been fixed in advance; or transposing these segments to another market. Also forbidden: presenting this synthetic case as an observed measurement.
16 · PROTOCOL
Possible marketing decision
The reasonable decision is to give each segment a distinct relationship programme, sized on its share and its value gap, then to measure the effect of those programmes with an experiment rather than by comparing segments with one another. The segment densest in value justifies the largest effort, but nothing yet guarantees it will respond best: the next campaign, with a control group inside each segment, will say.
17 · PROTOCOL
When to use or avoid the method
Use this method when the team knows how many programmes it can run and which variables describe what it wants to distinguish. Avoid it as a pure exploration tool: without a declared purpose, no criterion says that one split is better than another. Avoid it too when the chosen variables mostly measure the acquisition channel, when the base has just been renewed, or when the real objective is to measure the effect of an action, which calls for an experiment.
18 · PROTOCOL
Implementations and final deliverable
The CC0 CSV holds the nine hundred synthetic customers. Python and R are the reference implementations; SPSS carries the same computation in a Python block; SAS receives the same four declared starting points but applies its own assignment order and convergence rule, and checks balance and usefulness rather than reproducing every decimal. The final deliverable gathers the data, the variable dictionary, the declared variables and distance, the number of segments and its rationale, the declared start, the four thresholds, the seven numbered steps, the profiles, the four checks, the verdict and the software versions.
Synthetic data · CC0
msc-p031-segmentation-panel.csv ↓Reproducibility protocol
msc-p031-reproducibility-readme.md ↓Python reference · MIT
msc-p031-reference.py ↓R reference · MIT
msc-p031-reference.R ↓SPSS implementation · MIT
msc-p031-secondary.sps ↓SAS implementation · MIT
msc-p031-secondary.sas ↓19 · PROTOCOL
Sources and evidence level
Hennig recalls that a partitioning analysis becomes scientific through the transparency of its choices, not the uniqueness of its result, and that the notion of a true group remains rarely defined. Von Luxburg, Williamson and Guyon show that no general procedure ranks partitioning methods without taking the context of use into account. Romano et al. place the adjusted Rand index among chance-corrected measures. Ullmann, Hennig and Boulesteix formalize the validation of a split on other observations, by splitting the file into two halves. Founding references: Hennig (2015, arXiv preprint; published in Pattern Recognition Letters, peer-reviewed) and von Luxburg, Williamson and Guyon (2012, peer-reviewed proceedings, no DOI). Recent developments: Romano et al. (2016, Journal of Machine Learning Research, peer-reviewed, no DOI) and Ullmann, Hennig and Boulesteix (2021, arXiv preprint; published in WIREs Data Mining and Knowledge Discovery, peer-reviewed).
- Hennig (2015) Full text verified, with a short excerpt located in the source.
- von Luxburg, Williamson & Guyon (2012) Full text verified, with a short excerpt located in the source.
- Romano et al. (2016) Full text verified, with a short excerpt located in the source.
- Ullmann, Hennig & Boulesteix (2021) Full text verified, with a short excerpt located in the source.
Method connections

