How do you test segmentation stability?
Stability assesses whether a similar structure reappears under resampling, time variation or reasonable feature perturbation. It does not guarantee business usefulness.
Scientific editorial team: Marketing Science Center
Direct answer
Compare solutions and quantify assignment robustness.
Stability assesses whether a similar structure reappears under resampling, time variation or reasonable feature perturbation. It does not guarantee business usefulness.
01 · PROTOCOL
Operational summary
A four-segment solution is proposed on 360 customers. It is rebuilt on 40 resamples drawn with replacement, and each segment’s recovery is measured by the Jaccard index. The weakest segment is recovered at only 0.664749, below the declared threshold of 0.75. Worse: a structureless reference, obtained by permuting each feature, recovers its own segments at 0.829190, so better than the real file. Verdict: PROPOSED_SEGMENTATION_NOT_STABLE. The three-segment solution, by contrast, holds at 0.986949.
02 · PROTOCOL
Concrete marketing situation
A marketing department has received a four-group segmentation from an agency, complete with names, personas and a programme for each group. Before committing four separate budgets, the team asks a simple question: if the same work were redone on a slightly different sample of the same customers, would these four groups come back? Nobody is asking whether the groups are true. The question is whether they hold.
03 · PROTOCOL
Scientific question
For these 360 customers, these three declared features, this declared distance and this declared resampling procedure, does each of the four proposed segments reappear often enough to justify a budget of its own? And does that level of reappearance exceed what a file with no structure at all would already produce through resampling alone? The question is neither the existence of the groups nor their usefulness, but their reproducibility under perturbation.
04 · PROTOCOL
Why the simple approach can fail
Looking at the solution once says nothing: a partitioning algorithm always returns the number of groups requested, and on a single draw that number always looks clean. Rerunning the computation without changing the data says nothing either, since the start is deterministic. And measuring a high recovery with nothing to compare it to can mislead completely: on this file the structureless reference reaches 0.829190 at four segments, a figure one would happily read as proof of soundness if it were not set against something else.
05 · PROTOCOL
Method intuition
The idea fits in one sentence: perturb the file, redo exactly the same work, and see what comes back. Each perturbation is a resample drawn with replacement, the same size as the file. After each rebuild, all original customers are reassigned to the new centres, so that both splits cover the same people and compare without any trick. Each original segment then receives the best recovery it obtains against the rebuilt segments. The whole thing is measured a second time on a structureless reference, to learn what noise alone produces.
06 · PROTOCOL
Required data
Frozen before any reading: the customer identifier ordered without gaps, recency, frequency, average basket, the declared transformations, the standardization, the distance, the deterministic start, the numbers of segments examined (2, 3 and 4), the proposed number (4), the 40 resamples, the arithmetic recursion that produces them, the structureless reference and the two thresholds. The sealed file holds 360 customers. The case is synthetic and declared as such.
07 · PROTOCOL
Formal model and symbols
Each customer is a vector of three standardized coordinates: log recency, frequency, log basket. The distance is squared Euclidean and a segment gathers the customers closer to their centre than to any other. The recovery of a segment S by a rebuilt segment R is the Jaccard index, the number of customers in both divided by the number of customers in either. Resamples come from a declared recursion: the next state is (1103515245 × state + 12345) modulo 2^31, and the customer drawn is the state modulo the number of customers.
08 · PROTOCOL
Declared calculation
The declared computation runs seven steps: validate the schema and refuse any non-conforming file; transform and standardize the three features over the whole file; build the structureless reference by permuting each feature independently with the same recursion; for each number of segments examined, build the split on the whole file and record the composition of each segment; draw the 40 resamples, rebuild, reassign all original customers and give each segment its best Jaccard, failing if a resample cannot produce the requested number; average over the resamples, repeat identically on the structureless reference, then compute the weakest recovery, the reference one and their margin; apply the two thresholds, then the verdict rule to the proposed number.
09 · PROTOCOL
End-to-end numeric example
Synthetic illustration — teaching values, not observed
On the sealed file: 360 customers, 40 resamples. At two segments, recoveries 0.977481 and 0.976322; the weakest is 0.976322, the structureless reference 0.647971, the margin 0.328351, hence STABLE. At three segments: 0.990139, 0.986949 and 0.997222; weakest 0.986949, reference 0.606088, margin 0.380860, hence STABLE. At four segments, the proposed number: 0.848874, 0.664749, 0.827341 and 0.979621; weakest 0.664749, below the 0.75 threshold, reference 0.829190 and margin −0.164441, hence NOT_STABLE. Verdict: PROPOSED_SEGMENTATION_NOT_STABLE.
10 · PROTOCOL
Validity assumptions
The procedure assumes that the customers in the file can be treated as interchangeable observations from one population, which is what licenses drawing with replacement; that the three features are the ones describing what the team wants to distinguish; and that the reference obtained by permuting each feature does destroy the joint structure while preserving each marginal distribution. It does not assume that four segments exist, nor that three is the right number. Three things are not testable here: the relevance of the features, the persistence of the split over time, and a segment’s differentiated response to an action.
11 · PROTOCOL
Diagnostics and uncertainty
The mean recovery is taken over 40 resamples: no interval accompanies it, because at that number of draws the uncertainty on the mean would remain of the same order as the gaps one is trying to read, and the page says so rather than displaying a precision it does not have. The decisive diagnostic is not the raw level but the margin: at four segments it is negative, −0.164441, meaning the real file reproduces its segments less well than a file with no structure. None of these numbers is a p-value, and none tests a hypothesis.
12 · PROTOCOL
Robustness and alternatives
Alternatives declared before results: replace drawing with replacement by subsampling without replacement at a fixed size, and check that the conclusion does not depend on the scheme; replace the Jaccard index by the adjusted Rand index, which judges the whole split instead of each segment; raise the number of resamples to tighten the means; or perturb the features with a declared noise rather than by drawing. Each variant must be announced before reading and reported even if it overturns the verdict.
13 · PROTOCOL
Result interpretation
The proposed solution does not hold. One of its four segments dissolves as soon as the computation is replayed on slightly different customers, and the whole reproduces less well than a file with no structure. The three-segment split, by contrast, comes back almost intact. But the reading stops there: the two-segment split is stable as well, which shows that a good recovery does not designate the right number of groups. Stability eliminates candidates; it crowns none.
14 · PROTOCOL
Allowed conclusions
Allowed: refusing to commit four budgets to the proposed solution; publishing each segment’s recovery, the structureless reference and the margin; saying that one of the four segments dissolves under resampling; keeping three segments as a serious candidate provided it is justified by something other than stability; and requiring any agency to supply these three figures with every segmentation delivered. Also allowed: saying that this result holds for these features and this period, and redoing the examination if either changes.
15 · PROTOCOL
Forbidden conclusions
Forbidden: concluding that three is the true number of segments, when two also clears the thresholds; reading a high recovery as proof that a structure exists, when the structureless reference reaches 0.829190 at four segments; presenting stability as a measure of business usefulness; concluding that a stable segment will respond better to an offer; or rerunning the computation with several thresholds and publishing only the convenient one. Also forbidden: presenting this synthetic case as an observed measurement.
16 · PROTOCOL
Possible marketing decision
The reasonable decision is to send the four-group segmentation back to its author with the three figures that fault it, and to fund none of the four programmes as they stand. If the team wants to move forward, it restarts from the three-segment solution, justifies it by its capacity to run programmes and by the value gap between groups, then measures the real effect of each programme with an experiment carrying a control group inside each segment. Stability conditions the right to continue; it does not replace measurement.
17 · PROTOCOL
When to use or avoid the method
Use this examination whenever a segmentation is delivered as the basis for budgets, and before any production rollout. Avoid it as the sole criterion for choosing the number of segments, since it designates none. Avoid it too when the base has just been renewed, when the features do not describe what one wants to distinguish, or when the real question is the effect of an action on a segment, which calls for an experiment rather than a resampling.
18 · PROTOCOL
Implementations and final deliverable
The CC0 CSV holds the 360 synthetic customers. Python and R are the reference implementations; SPSS carries the same computation in a Python block; SAS reproduces the declared recursion exactly, so it draws the same resamples, but delegates the grouping to its own procedure and checks the flags rather than every decimal. The final deliverable gathers the data, the variable dictionary, the declared features and distance, the resampling recursion, the structureless reference, the two thresholds, the seven numbered steps, each segment’s recovery for each number examined, the margins, the flags, the verdict and the software versions.
Synthetic data · CC0
msc-p032-stability-panel.csv ↓Reproducibility protocol
msc-p032-reproducibility-readme.md ↓Python reference · MIT
msc-p032-reference.py ↓R reference · MIT
msc-p032-reference.R ↓SPSS implementation · MIT
msc-p032-secondary.sps ↓SAS implementation · MIT
msc-p032-secondary.sas ↓19 · PROTOCOL
Sources and evidence level
Von Luxburg explains that testing stability requires rerunning the algorithm on slightly different data sets, and that a stable result at too small a number of segments supports no useful conclusion; she also recalls that the score is compared with one obtained from a structureless reference distribution, and that drawing with replacement is the standard scheme of the bootstrap literature. Hennig notes that stability is easier to reach with fewer groups, and that features must be chosen for the question asked. Ullmann, Hennig and Boulesteix place stability analysis inside a validation framework and recall the common practice of keeping the number of segments that is most stable. Ullmann and co-authors finally show that the number of segments is a setting whose consequences must be examined, and that guarding against over-optimism demands rules fixed in advance. Founding references: von Luxburg (2010, arXiv preprint; published in Foundations and Trends in Machine Learning, peer-reviewed) and Hennig (2015, arXiv preprint; published in Pattern Recognition Letters, peer-reviewed). Recent developments: Ullmann, Hennig and Boulesteix (2021, arXiv preprint; published in WIREs Data Mining and Knowledge Discovery, peer-reviewed) and Ullmann et al. (2022, Advances in Data Analysis and Classification, CC BY open access, peer-reviewed).
- von Luxburg (2010) Full text verified, with a short excerpt located in the source.
- Hennig (2015) Full text verified, with a short excerpt located in the source.
- Ullmann, Hennig & Boulesteix (2021) Full text verified, with a short excerpt located in the source.
- Ullmann et al. (2022) Full text verified, with a short excerpt located in the source.
Method connections

