Research & Evidence

Search for a method

Search titles, questions, territories and MSC identifiers.

41 results
  1. MSC-P-001How do you turn a marketing claim into a testable question?Decision Science↗
  2. MSC-P-002Correlation or causality: what can an analysis actually support?Marketing Measurement↗
  3. MSC-P-003How should uncertainty in a marketing result be expressed?Decision Science↗
  4. MSC-P-004Statistical significance or effect size: which result should be interpreted?Decision Science↗
  5. MSC-P-005How do you measure a marketing construct that is not directly observable?Market Research↗
  6. MSC-P-006How do you design and validate a measurement scale?Market Research↗
  7. MSC-P-007Alpha or omega: how should scale reliability be assessed?Market Research↗
  8. MSC-P-009PCA, EFA or CFA: which method should you choose?Market Research↗
  9. MSC-P-010When should you run a marketing experiment?Marketing Measurement↗
  10. MSC-P-011How do you design an A/B test that actually estimates an effect?Marketing Measurement↗
  11. MSC-P-012How many observations does an experiment need?Decision Science↗
  12. MSC-P-013How do you measure campaign incrementality with a control group?Marketing Measurement↗
  13. MSC-P-017How do you detect selection, contamination and attrition in an experiment?Marketing Measurement↗
  14. MSC-P-018Predictive or causal regression: what are you trying to estimate?Marketing Models↗
  15. MSC-P-019How do you diagnose a marketing regression before interpreting it?Marketing Models↗
  16. MSC-P-022How do you estimate price elasticity and its uncertainty?Pricing Science↗
  17. MSC-P-026Logit vs Probit: how do you choose for purchase probability?Customer Science↗
  18. MSC-P-029Which customers have the highest probability of churn?Customer Science↗
  19. MSC-P-027TAM, UTAUT or UTAUT2: which framework should be used to study technology acceptance?Market Research↗
  20. MSC-H-001Measurement and causality: how can a marketing effect be established?Marketing Measurement↗
  21. MSC-H-002Marketing response models: shape, delay and saturationMarketing Models↗
  22. MSC-H-003Pricing science: connecting price, demand and contributionPricing Science↗
  23. MSC-H-004Customer and choice science: behavior, value and heterogeneityCustomer Science↗
  24. MSC-H-005Measurement science: building valid indicatorsMarket Research↗
  25. MSC-H-006Statistical decision methods: choose, quantify, validateDecision Science↗
  26. MSC-P-008How do you validate a marketing measurement scale?Market Research↗
  27. MSC-P-014How do you design a marketing geo experiment?Marketing Measurement↗
  28. MSC-P-015How do you estimate an effect with difference-in-differences?Marketing Measurement↗
  29. MSC-P-020How do you address price endogeneity?Pricing Science↗
  30. MSC-P-021Fixed or random effects: which panel model should you choose?Marketing Models↗
  31. MSC-P-023How do you estimate a demand function?Pricing Science↗
  32. MSC-P-024How do you simulate a price-volume-margin scenario?Pricing Science↗
  33. MSC-P-028How do you estimate CLV with BG/NBD and Gamma-Gamma?Customer Science↗
  34. MSC-P-030How do you analyze retention with a survival model?Customer Science↗
  35. MSC-P-043What is a subscriber worth with only six months of retention data?Customer Science↗
  36. MSC-P-031How do you build a useful customer segmentation?Customer Science↗
  37. MSC-P-032How do you test segmentation stability?Customer Science↗
  38. MSC-P-033How do you validate a marketing forecast?Decision Science↗
  39. MSC-P-034How do you build a Monte Carlo simulation for a marketing decision?Decision Science↗
  40. MSC-P-035How do you model saturation and adstock?Marketing Models↗
  41. MSC-P-039Which statistical test should you choose?Decision Science↗
← All methods
METHOD DOSSIERMSC-P-032Customer scienceVerified scientific dossier

How do you test segmentation stability?

Stability assesses whether a similar structure reappears under resampling, time variation or reasonable feature perturbation. It does not guarantee business usefulness.

Scientific editorial team: Marketing Science Center

Direct answer

Compare solutions and quantify assignment robustness.

Stability assesses whether a similar structure reappears under resampling, time variation or reasonable feature perturbation. It does not guarantee business usefulness.

Hennig, 2007Rousseeuw, 1987

01 · PROTOCOL

Operational summary

A four-segment solution is proposed on 360 customers. It is rebuilt on 40 resamples drawn with replacement, and each segment’s recovery is measured by the Jaccard index. The weakest segment is recovered at only 0.664749, below the declared threshold of 0.75. Worse: a structureless reference, obtained by permuting each feature, recovers its own segments at 0.829190, so better than the real file. Verdict: PROPOSED_SEGMENTATION_NOT_STABLE. The three-segment solution, by contrast, holds at 0.986949.

02 · PROTOCOL

Concrete marketing situation

A marketing department has received a four-group segmentation from an agency, complete with names, personas and a programme for each group. Before committing four separate budgets, the team asks a simple question: if the same work were redone on a slightly different sample of the same customers, would these four groups come back? Nobody is asking whether the groups are true. The question is whether they hold.

03 · PROTOCOL

Scientific question

For these 360 customers, these three declared features, this declared distance and this declared resampling procedure, does each of the four proposed segments reappear often enough to justify a budget of its own? And does that level of reappearance exceed what a file with no structure at all would already produce through resampling alone? The question is neither the existence of the groups nor their usefulness, but their reproducibility under perturbation.

04 · PROTOCOL

Why the simple approach can fail

Looking at the solution once says nothing: a partitioning algorithm always returns the number of groups requested, and on a single draw that number always looks clean. Rerunning the computation without changing the data says nothing either, since the start is deterministic. And measuring a high recovery with nothing to compare it to can mislead completely: on this file the structureless reference reaches 0.829190 at four segments, a figure one would happily read as proof of soundness if it were not set against something else.

05 · PROTOCOL

Method intuition

The idea fits in one sentence: perturb the file, redo exactly the same work, and see what comes back. Each perturbation is a resample drawn with replacement, the same size as the file. After each rebuild, all original customers are reassigned to the new centres, so that both splits cover the same people and compare without any trick. Each original segment then receives the best recovery it obtains against the rebuilt segments. The whole thing is measured a second time on a structureless reference, to learn what noise alone produces.

06 · PROTOCOL

Required data

Frozen before any reading: the customer identifier ordered without gaps, recency, frequency, average basket, the declared transformations, the standardization, the distance, the deterministic start, the numbers of segments examined (2, 3 and 4), the proposed number (4), the 40 resamples, the arithmetic recursion that produces them, the structureless reference and the two thresholds. The sealed file holds 360 customers. The case is synthetic and declared as such.

07 · PROTOCOL

Formal model and symbols

Each customer is a vector of three standardized coordinates: log recency, frequency, log basket. The distance is squared Euclidean and a segment gathers the customers closer to their centre than to any other. The recovery of a segment S by a rebuilt segment R is the Jaccard index, the number of customers in both divided by the number of customers in either. Resamples come from a declared recursion: the next state is (1103515245 × state + 12345) modulo 2^31, and the customer drawn is the state modulo the number of customers.

08 · PROTOCOL

Declared calculation

The declared computation runs seven steps: validate the schema and refuse any non-conforming file; transform and standardize the three features over the whole file; build the structureless reference by permuting each feature independently with the same recursion; for each number of segments examined, build the split on the whole file and record the composition of each segment; draw the 40 resamples, rebuild, reassign all original customers and give each segment its best Jaccard, failing if a resample cannot produce the requested number; average over the resamples, repeat identically on the structureless reference, then compute the weakest recovery, the reference one and their margin; apply the two thresholds, then the verdict rule to the proposed number.

09 · PROTOCOL

End-to-end numeric example

Synthetic illustration — teaching values, not observed

On the sealed file: 360 customers, 40 resamples. At two segments, recoveries 0.977481 and 0.976322; the weakest is 0.976322, the structureless reference 0.647971, the margin 0.328351, hence STABLE. At three segments: 0.990139, 0.986949 and 0.997222; weakest 0.986949, reference 0.606088, margin 0.380860, hence STABLE. At four segments, the proposed number: 0.848874, 0.664749, 0.827341 and 0.979621; weakest 0.664749, below the 0.75 threshold, reference 0.829190 and margin −0.164441, hence NOT_STABLE. Verdict: PROPOSED_SEGMENTATION_NOT_STABLE.

10 · PROTOCOL

Validity assumptions

The procedure assumes that the customers in the file can be treated as interchangeable observations from one population, which is what licenses drawing with replacement; that the three features are the ones describing what the team wants to distinguish; and that the reference obtained by permuting each feature does destroy the joint structure while preserving each marginal distribution. It does not assume that four segments exist, nor that three is the right number. Three things are not testable here: the relevance of the features, the persistence of the split over time, and a segment’s differentiated response to an action.

11 · PROTOCOL

Diagnostics and uncertainty

The mean recovery is taken over 40 resamples: no interval accompanies it, because at that number of draws the uncertainty on the mean would remain of the same order as the gaps one is trying to read, and the page says so rather than displaying a precision it does not have. The decisive diagnostic is not the raw level but the margin: at four segments it is negative, −0.164441, meaning the real file reproduces its segments less well than a file with no structure. None of these numbers is a p-value, and none tests a hypothesis.

12 · PROTOCOL

Robustness and alternatives

Alternatives declared before results: replace drawing with replacement by subsampling without replacement at a fixed size, and check that the conclusion does not depend on the scheme; replace the Jaccard index by the adjusted Rand index, which judges the whole split instead of each segment; raise the number of resamples to tighten the means; or perturb the features with a declared noise rather than by drawing. Each variant must be announced before reading and reported even if it overturns the verdict.

13 · PROTOCOL

Result interpretation

The proposed solution does not hold. One of its four segments dissolves as soon as the computation is replayed on slightly different customers, and the whole reproduces less well than a file with no structure. The three-segment split, by contrast, comes back almost intact. But the reading stops there: the two-segment split is stable as well, which shows that a good recovery does not designate the right number of groups. Stability eliminates candidates; it crowns none.

14 · PROTOCOL

Allowed conclusions

Allowed: refusing to commit four budgets to the proposed solution; publishing each segment’s recovery, the structureless reference and the margin; saying that one of the four segments dissolves under resampling; keeping three segments as a serious candidate provided it is justified by something other than stability; and requiring any agency to supply these three figures with every segmentation delivered. Also allowed: saying that this result holds for these features and this period, and redoing the examination if either changes.

15 · PROTOCOL

Forbidden conclusions

Forbidden: concluding that three is the true number of segments, when two also clears the thresholds; reading a high recovery as proof that a structure exists, when the structureless reference reaches 0.829190 at four segments; presenting stability as a measure of business usefulness; concluding that a stable segment will respond better to an offer; or rerunning the computation with several thresholds and publishing only the convenient one. Also forbidden: presenting this synthetic case as an observed measurement.

16 · PROTOCOL

Possible marketing decision

The reasonable decision is to send the four-group segmentation back to its author with the three figures that fault it, and to fund none of the four programmes as they stand. If the team wants to move forward, it restarts from the three-segment solution, justifies it by its capacity to run programmes and by the value gap between groups, then measures the real effect of each programme with an experiment carrying a control group inside each segment. Stability conditions the right to continue; it does not replace measurement.

17 · PROTOCOL

When to use or avoid the method

Use this examination whenever a segmentation is delivered as the basis for budgets, and before any production rollout. Avoid it as the sole criterion for choosing the number of segments, since it designates none. Avoid it too when the base has just been renewed, when the features do not describe what one wants to distinguish, or when the real question is the effect of an action on a segment, which calls for an experiment rather than a resampling.

18 · PROTOCOL

Implementations and final deliverable

The CC0 CSV holds the 360 synthetic customers. Python and R are the reference implementations; SPSS carries the same computation in a Python block; SAS reproduces the declared recursion exactly, so it draws the same resamples, but delegates the grouping to its own procedure and checks the flags rather than every decimal. The final deliverable gathers the data, the variable dictionary, the declared features and distance, the resampling recursion, the structureless reference, the two thresholds, the seven numbered steps, each segment’s recovery for each number examined, the margins, the flags, the verdict and the software versions.

19 · PROTOCOL

Sources and evidence level

Von Luxburg explains that testing stability requires rerunning the algorithm on slightly different data sets, and that a stable result at too small a number of segments supports no useful conclusion; she also recalls that the score is compared with one obtained from a structureless reference distribution, and that drawing with replacement is the standard scheme of the bootstrap literature. Hennig notes that stability is easier to reach with fewer groups, and that features must be chosen for the question asked. Ullmann, Hennig and Boulesteix place stability analysis inside a validation framework and recall the common practice of keeping the number of segments that is most stable. Ullmann and co-authors finally show that the number of segments is a setting whose consequences must be examined, and that guarding against over-optimism demands rules fixed in advance. Founding references: von Luxburg (2010, arXiv preprint; published in Foundations and Trends in Machine Learning, peer-reviewed) and Hennig (2015, arXiv preprint; published in Pattern Recognition Letters, peer-reviewed). Recent developments: Ullmann, Hennig and Boulesteix (2021, arXiv preprint; published in WIREs Data Mining and Knowledge Discovery, peer-reviewed) and Ullmann et al. (2022, Advances in Data Analysis and Classification, CC BY open access, peer-reviewed).

  1. von Luxburg (2010) Full text verified, with a short excerpt located in the source.
  2. Hennig (2015) Full text verified, with a short excerpt located in the source.
  3. Ullmann, Hennig & Boulesteix (2021) Full text verified, with a short excerpt located in the source.
  4. Ullmann et al. (2022) Full text verified, with a short excerpt located in the source.

Method connections

Parent territoryCustomer and choice science: behavior, value and heterogeneityValidatesHow do you build a useful customer segmentation?Compare withHow do you validate a marketing forecast?

Read next

MSC-H-004Customer and choice science: behavior, value and heterogeneity→MSC-P-031How do you build a useful customer segmentation?→MSC-P-019How do you diagnose a marketing regression before interpreting it?→MSC-P-033How do you validate a marketing forecast?→