Research & Evidence

Search for a method

Search titles, questions, territories and MSC identifiers.

41 results
  1. MSC-P-001How do you turn a marketing claim into a testable question?Decision Science↗
  2. MSC-P-002Correlation or causality: what can an analysis actually support?Marketing Measurement↗
  3. MSC-P-003How should uncertainty in a marketing result be expressed?Decision Science↗
  4. MSC-P-004Statistical significance or effect size: which result should be interpreted?Decision Science↗
  5. MSC-P-005How do you measure a marketing construct that is not directly observable?Market Research↗
  6. MSC-P-006How do you design and validate a measurement scale?Market Research↗
  7. MSC-P-007Alpha or omega: how should scale reliability be assessed?Market Research↗
  8. MSC-P-009PCA, EFA or CFA: which method should you choose?Market Research↗
  9. MSC-P-010When should you run a marketing experiment?Marketing Measurement↗
  10. MSC-P-011How do you design an A/B test that actually estimates an effect?Marketing Measurement↗
  11. MSC-P-012How many observations does an experiment need?Decision Science↗
  12. MSC-P-013How do you measure campaign incrementality with a control group?Marketing Measurement↗
  13. MSC-P-017How do you detect selection, contamination and attrition in an experiment?Marketing Measurement↗
  14. MSC-P-018Predictive or causal regression: what are you trying to estimate?Marketing Models↗
  15. MSC-P-019How do you diagnose a marketing regression before interpreting it?Marketing Models↗
  16. MSC-P-022How do you estimate price elasticity and its uncertainty?Pricing Science↗
  17. MSC-P-026Logit vs Probit: how do you choose for purchase probability?Customer Science↗
  18. MSC-P-029Which customers have the highest probability of churn?Customer Science↗
  19. MSC-P-027TAM, UTAUT or UTAUT2: which framework should be used to study technology acceptance?Market Research↗
  20. MSC-H-001Measurement and causality: how can a marketing effect be established?Marketing Measurement↗
  21. MSC-H-002Marketing response models: shape, delay and saturationMarketing Models↗
  22. MSC-H-003Pricing science: connecting price, demand and contributionPricing Science↗
  23. MSC-H-004Customer and choice science: behavior, value and heterogeneityCustomer Science↗
  24. MSC-H-005Measurement science: building valid indicatorsMarket Research↗
  25. MSC-H-006Statistical decision methods: choose, quantify, validateDecision Science↗
  26. MSC-P-008How do you validate a marketing measurement scale?Market Research↗
  27. MSC-P-014How do you design a marketing geo experiment?Marketing Measurement↗
  28. MSC-P-015How do you estimate an effect with difference-in-differences?Marketing Measurement↗
  29. MSC-P-020How do you address price endogeneity?Pricing Science↗
  30. MSC-P-021Fixed or random effects: which panel model should you choose?Marketing Models↗
  31. MSC-P-023How do you estimate a demand function?Pricing Science↗
  32. MSC-P-024How do you simulate a price-volume-margin scenario?Pricing Science↗
  33. MSC-P-028How do you estimate CLV with BG/NBD and Gamma-Gamma?Customer Science↗
  34. MSC-P-030How do you analyze retention with a survival model?Customer Science↗
  35. MSC-P-043What is a subscriber worth with only six months of retention data?Customer Science↗
  36. MSC-P-031How do you build a useful customer segmentation?Customer Science↗
  37. MSC-P-032How do you test segmentation stability?Customer Science↗
  38. MSC-P-033How do you validate a marketing forecast?Decision Science↗
  39. MSC-P-034How do you build a Monte Carlo simulation for a marketing decision?Decision Science↗
  40. MSC-P-035How do you model saturation and adstock?Marketing Models↗
  41. MSC-P-039Which statistical test should you choose?Decision Science↗
← All methods
METHOD DOSSIERMSC-P-043Customer scienceVerified scientific dossier

What is a subscriber worth with only six months of retention data?

Subscriber retention rises with tenure, so a constant rate undervalues your base. The sBG model projects retention and subscriber CLV.

Scientific editorial team: Marketing Science Center

Direct answer

Project a subscriber cohort’s retention and estimate the discounted value of a new subscriber and of the remaining subscribers, with an interval.

Subscriber retention rises with tenure, so a constant rate undervalues your base. The sBG model projects retention and subscriber CLV.

Fader & Hardie, 2007Fader & Hardie, 2010

01

In short

To estimate what a subscriber is worth, the textbook method takes a retention rate and assumes it stays constant [C19]. Yet within a cohort, this rate rises with tenure: the most volatile subscribers leave first, and those who remain are, on average, more loyal [C09]. A constant rate ignores this sorting and underestimates the value of the subscriber base [C20]. Fader and Hardie's shifted-beta-geometric (sBG) model describes this sorting with two parameters and projects retention beyond the observed history.

On synthetic data, where the truth is known, a cohort of 2,000 monthly subscribers is observed for six months: the constant rate puts the value of a new subscriber at €103, the sBG model at €290, against a true value of €279. Across 100 independent draws, the constant rate underestimates the value of the remaining subscribers in 100 cases out of 100. Two caveats weigh on these figures: more than half of the value lies beyond month 24, which the history cannot verify, and the result depends heavily on the chosen discount rate.

02

The situation

A monthly subscription service — a subscription box, software, an online newspaper — acquired 2,000 subscribers in January. At the start of each month, the subscriber pays and the company earns a margin of €12; at the end of the month, the subscriber renews or cancels. Six months later, 1,058 subscribers are still there. The finance department asks for two figures: how much a new subscriber is worth, to set the maximum acquisition cost, and how much the 1,058 remaining subscribers are worth, to know what the base to be retained represents.

The data on this page are synthetic: we know that the probability of cancelling varies from one subscriber to another according to a beta distribution with parameters 0.5 and 2.5 — a distribution that describes how this probability is spread across subscribers — which gives an average churn of 1 in 6 per month, a high level for a subscription.

03

The scientific question

Given the cancellations observed during the first six months of a cohort, what is the expected present value of a new subscriber's future margins, and that of a subscriber still active after six months? Both values require projecting the cohort's survival curve well beyond the observed horizon: summing only the observed part of the curve underestimates lifetime and value [C02].

04

Why the simple method fails

The observed retention rates rise markedly over the period: 83.4%, 87.8%, 89.3%, 92.5%, 93.9% and 93.2%. Two common shortcuts summarize them in a single figure:

  • the overall rate, weighted by renewal points: 942 cancellations out of 8,781 renewal points, i.e. a constant retention rate of 89.3%;
  • the last observed rate, seen as more representative of today's cohort: 93.2%.

In both cases, every future month is assumed to look like the past. Yet the rise in rates does not come from each subscriber becoming more loyal over time: in our data, the individual probability is constant by construction. It comes from sorting: subscribers with a high probability of cancelling leave early, and the cohort becomes concentrated on the loyal ones [C08, C09]. This phenomenon is called "the ruse of heterogeneity" [C10]. Fitting a straight line, a parabola or an exponential to the observed curve does not help: these curves hug the past data and go off the rails once they are extended [C03].

05

The intuition

The sBG model tells a simple story [C04]:

  1. at the end of each month, each subscriber "tosses a coin": they cancel or they renew;
  2. for a given subscriber, the probability of cancelling, denoted θ, does not change over time;
  3. this probability varies from one subscriber to another [C05], according to a beta distribution [C06], a flexible distribution bounded between 0 and 1.

Two parameters, α and β, summarize the whole cohort. They are estimated on the observed months, and the retention of each future month is then calculated. The model mechanically produces retention rates that rise with tenure, even though each subscriber keeps the same behavior [C08].

06

The data you need

For each acquisition cohort, the number of subscribers still active at each renewal point. Aggregate counts are enough: the model can even be estimated on the share of the cohort still active [C31] — but the actual counts are needed to measure uncertainty. You also need the margin per period, collected at the start of the period, and a discount rate expressed in the same time unit as the renewal points: here 1% per month, i.e. about 12.7% per year.

The method assumes fixed renewal points (monthly, annual) and cancellations that the company observes [C01]. If cancellation can happen at any moment, the continuous-time equivalent is needed [C15]; if the company does not see the customer leave, as in non-subscription retail, other models are needed [C16].

07

The formal model

This section is for practitioners. Decision-makers can skip it and resume at "The declared calculation": the "In plain words" sentence at its end is enough.

  • T: the number of monthly payments made by a subscriber, i.e. their lifetime in months; T ≥ 1.
  • θ: the probability, specific to each subscriber, of cancelling at a renewal point; it is constant over time. If θ were known, we would have S(t | θ) = (1 − θ)ᵗ.
  • θ follows a beta distribution with parameters α > 0 and β > 0 [C06]. Its mean α / (α + β) is the average probability of cancelling; the polarization index φ = 1 / (α + β + 1) measures heterogeneity [C34]: near 0, all subscribers are alike; near 1, they split into the very loyal and the very volatile.
  • S(t) = P(T > t): the share of the cohort still active after t renewal points; S(0) = 1.
  • Retention rate at renewal point t: rₜ = S(t) / S(t − 1) = (β + t − 1) / (α + β + t − 1), hence S(t) = r₁ × r₂ × … × rₜ. The probabilities are computed recursively, without the beta function [C07]: P(T = 1) = α / (α + β) and P(T = t) = P(T = t − 1) × (β + t − 2) / (α + β + t − 1).
  • Log-likelihood over six renewal points, with nₜ cancellations at renewal point t, s₆ subscribers still active and ln the natural logarithm: LL(α, β) = Σₜ₌₁⁶ nₜ ln P(T = t) + s₆ ln S(6) [C38]. We look for the values of α and β that maximize it [C11].
  • m: margin per month, collected at the start of the month; d: monthly discount rate.
  • Value of a new subscriber, or CLV (customer lifetime value): CLV = m × Σₜ₌₀^∞ S(t) / (1 + d)ᵗ [C33].
  • Residual value of a subscriber still active after n renewal points, measured at the moment they have just paid their (n + 1)th monthly payment: RV(n) = m × Σₜ₌ₙ₊₁^∞ [S(t) / S(n)] / (1 + d)ᵗ⁻ⁿ. The ratio S(t) / S(n) takes into account the tenure already accrued: a subscriber who has already renewed six times probably has a lower θ than average [C21]. Fader and Hardie measure this value just before the renewal decision and include the next payment in it; this choice of date is a convention [C39], and each formulation can be derived from the other.
  • Constant-rate (geometric) model, the limit of the sBG as polarization tends to 0 [C35]: θ̂ = cancellations / renewal points, S(t) = (1 − θ̂)ᵗ, hence CLV = m (1 + d) / (d + θ̂).

In plain words: each subscriber keeps their own probability of leaving; because the most volatile leave first, the cohort becomes more loyal, and a subscriber's value depends on how long they have already stayed.

08

The declared calculation

  1. Synthetic cohort generated by a declared linear congruential recurrence (seed 20261002; x ← 1103515245 · x + 12345 mod 2³¹), identical in Python and in R. Each subscriber's lifetime is drawn by inverting the true sBG survival function (α = 0.5, β = 2.5).
  2. Calibration on the first six renewal points; months 7 to 24 are held out to judge the projection.
  3. Constant-rate model: maximum likelihood estimation, which reduces to the ratio of cancellations to renewal points.
  4. sBG model: maximum likelihood on (ln α, ln β) with the Nelder-Mead simplex algorithm, written by hand; the program stops with an error if the algorithm does not converge. Standard errors from the numerical Hessian matrix on the log scale, carried back to α and β by the delta method: standard error of α = α × standard error of ln α.
  5. Last-rate variant: the last observed retention rate (month 6), extended unchanged.
  6. Values discounted at 1% per month (about 12.7% per year), margin of €12 per month, series truncated at 3,000 terms — the remainder is less than one cent.
  7. Uncertainty: parametric bootstrap, 200 cohorts simulated from the estimated parameters, each re-estimated; the interval runs from the 5th to the 195th of the 200 sorted values, i.e. an approximately 95% confidence interval.
  8. The whole procedure is repeated on 100 independent cohorts, to separate what is systematic from what is due to the luck of a single draw.
  9. Sensitivity: share of the value lying beyond months 24, 60 and 120; values limited to the first 24, 36 and 60 monthly payments; values for a discount rate of 0.5%, 1%, 1.5% and 2% per month, and of 10% per month, the authors' rate per renewal point, their renewal points being annual.
  10. Code check before use: both programs recover the estimates published by Fader and Hardie (2007) on their two datasets (α = 0.668 and β = 3.806; α = 0.704 and β = 1.182) [C12, C13], as well as the ten discounted residual lifetimes in Table 4 of Fader and Hardie (2010) [C32], and stop with an error if they do not recover them. The text of the 2007 article prints 0.688 for α on page 9; it is the appendix value, 0.668, that maximum likelihood recovers.

09

The complete worked example

The estimates. The constant-rate model gives churn of 10.7% per month. The sBG model gives α = 0.475 (standard error 0.048) and β = 2.371 (0.327), i.e. an average churn of 16.7% and a polarization index of 0.26 — close to the true values, 16.7% and 0.25. The sBG fits clearly better: log-likelihood −2,926.8 versus −2,992.4, and an Akaike information criterion — which trades off goodness of fit against the number of parameters, lower being better — lower by 129 points.

The projection, judged on the held-out months.

Share of the cohort still activeObservedConstant rateLast ratesBGTrue
Month 652.9%50.6%52.9%52.9%52.4%
Month 1239.9%25.6%34.7%40.7%39.9%
Month 1832.8%13.0%22.8%34.4%33.4%
Month 2429.0%6.6%14.9%30.4%29.4%
Mean absolute error, months 7 to 24—16.4 percentage points7.8 percentage points1.1 percentage points—

The sBG model estimates a retention rate of 93.9% in month 6, within the calibration period, then projects 96.6% in month 12 and 98.2% in month 24: the cohort settles on its loyal subscribers.

The values.

Present valueConstant rateLast ratesBGTrue
New subscriber€103.35—€290.26€279.26
Subscriber still active after 6 months€91.35€143.70€466.98€449.80
The 1,058 remaining subscribers€96,643€152,036€494,069€475,887

Compared with the true value, the constant rate underestimates the value of a new subscriber by 63% and that of the remaining subscribers by 80%; the last rate, by 68%. The sBG overestimates both values by about 4% on this draw.

What the history does not verify. Of the sBG value of a new subscriber, 56.8% lies beyond month 24, 32.4% beyond month 60 and 14.6% beyond month 120; for a remaining subscriber, these shares are 70.8%, 40.5% and 18.3%. Limited to the first monthly payments, the value of a new subscriber is as follows:

Value of a new subscriber, limited to the…Constant ratesBGTrue
first 24 monthly payments€98.00€125.41€123.73
first 36 monthly payments€102.13€155.38€152.60
first 60 monthly payments€103.28€196.12€191.46

Sensitivity to the discount rate.

Monthly rateNew subscriber, constant rateNew subscriber, sBGNew subscriber, trueRemaining subscriber, sBGRemaining subscriber, true
0.5%€107.41€431.17€409.83€726.32€692.07
1%€103.35€290.26€279.26€466.98€449.80
1.5%€99.61€229.24€221.98€356.37€345.21
2%€96.17€193.47€188.15€292.36€284.27

The choice of rate alone, between 0.5% and 2% per month, moves the sBG value of a new subscriber from €431 to €193: far more than the estimation uncertainty.

Across 100 independent cohorts.

EstimatorMeanStandard deviationBias
New subscriber, constant rate€102.13€2.56−€177.13
New subscriber, sBG€279.54€19.20+€0.28
Remaining subscriber, constant rate€90.13€2.56−€359.67
Remaining subscriber, last rate€156.56€15.96−€293.24
Remaining subscriber, sBG€449.86€33.50+€0.06

The bias is the mean of the 100 estimates minus the true value. For the sBG, it cannot be distinguished from zero: its simulation error is €1.92 for a new subscriber and €3.35 for a remaining subscriber. The constant rate and the last rate underestimate the value of the remaining subscribers in 100 draws out of 100. The constant-rate estimates vary little from one draw to the next: they are precise, but wrong.

10

Validity assumptions

  • Constant individual probability. The model attributes the entire rise in rates to sorting. Another possible explanation is that each subscriber genuinely changes with tenure [C28]; with aggregate data, the two explanations cannot be told apart, so the assumption cannot be tested here. Do your subscribers grow attached to the service over time, or do they tire of it? The discrete beta-Weibull model relaxes this assumption [C17].
  • No planned break. In monthly data, seasonality and the end of introductory offers create churn spikes; the sBG can then be replaced by a duration model with covariates [C29]. Does your €1 trial offer end in month 3?
  • Fixed renewal points and observed departure: the method applies neither to cancellations that can happen at any moment [C15], nor to non-subscription retail [C16]. Can your subscribers cancel mid-month, and do you always know that they have left?
  • A cohort representative of those to come. Cohorts acquired at different times or through different channels may behave differently [C37]. With several cohorts, the natural starting point is to pool them under common parameters; they can also be estimated separately or given their own parameters [C18]. Were your January and September cohorts acquired through the same channels, at the same price?
  • Constant margin per month: a subscriber who moves up to a higher plan requires modeling the margin separately. Does your margin per subscriber change with tenure?
  • Subscribers cancel independently of one another. Did a price change or an outage trigger a wave of simultaneous departures?

11

Diagnostics and uncertainty

  • Out-of-sample validation: on real data, keep the last months aside, estimate on the first ones and compare the projection with what was observed. Here, the sBG's mean error is 1.1 percentage points over months 7 to 24. This check covers only the observed horizon: it does not verify the 56.8% of a new subscriber's value that lies beyond month 24, which is determined by the shape of the beta distribution near θ = 0, whereas six months of history say little about it.
  • Comparing the fits: the likelihood ratio is 131.3. The constant-rate model is the limit of the sBG as polarization tends to 0 [C35], a value on the boundary of the parameter space: the usual χ² distribution does not apply exactly, and the figure should be read as an indicator, not as a test.
  • Estimation uncertainty: by parametric bootstrap, the approximately 95% confidence interval for the expected value of a new subscriber runs from €250 to €331, and that for the expected value of a remaining subscriber from €397 to €537. These intervals concern the expectation, not the value a given subscriber will actually generate; they assume the sBG model is correct and cover neither model error nor uncertainty about the future margin or the discount rate.
  • Standard errors of the parameters: 0.048 for α, 0.327 for β. With six months of history, β remains the least well-known parameter.
  • Discount rate: from 0.5% to 2% per month, the sBG value of a new subscriber ranges from €431 to €193; this choice, which belongs to the finance department, weighs more than the estimation uncertainty.
  • What the simulation proves, and what it does not: the data are generated by an sBG model, so the sBG is correct here by construction. The simulation measures the cost of ignoring heterogeneity when it exists; it does not show that the sBG suits your data.

12

Interpretation

The cohort's retention rate rises because the cohort changes composition, not because each subscriber changes. A constant rate treats month 25 like month 2, which amounts to making loyal subscribers leave too quickly. The error weighs most heavily on the value of the remaining subscribers: they are precisely the loyal ones. The model's authors find, through numerical analysis, an underestimation of the order of 25% to 50% in common situations [C26], and of 28% and 48% when they apply the parameters estimated on two real datasets to a hypothetical base acquiring 10,000 customers per year [C25]; they do not prove it formally [C30]. Our gap is larger, and the main reason is the per-period discount rate: at 1% per month, distant months, when only the loyal subscribers remain, weigh much more than at 10% per renewal point, the authors' convention (their renewal points are annual). Applied to our monthly renewal points and to our single cohort, not to the authors' five-cohort base, a rate of 10% per month brings the gap on the remaining subscribers down to −39% for the constant rate and −22% for the last rate. The more alike the subscribers, the smaller the constant rate's error [C23]; when heterogeneity is strong, it remains acceptable only if average churn is very low [C24].

13

Allowed and forbidden conclusions

  • Allowed: "On this cohort, if the rise in rates comes from the sorting of subscribers and at a discount rate of 1% per month, the expected value of a new subscriber is about €290 of discounted margin, with an approximately 95% confidence interval of €250 to €331."
  • Allowed: "In this simulation (sBG, α = 0.5, β = 2.5), valuing the base with a constant retention rate strongly underestimates it; this is the case in 100 draws out of 100."
  • Forbidden: concluding that subscribers become more loyal as they age. Aggregate data cannot distinguish sorting from genuine individual change [C28].
  • Forbidden: presenting the projected value as certain. A CLV is an expectation [C22], conditional on the model, the margin and the discount rate, and more than half of it is projected beyond the history.
  • Forbidden: applying the model to non-subscription customers, or to a cohort marked by the end of an introductory offer, without adapting it [C16, C29].
  • Forbidden: presenting these figures as a market result; the data are synthetic.

14

Possible marketing decision

  • Acquisition ceiling. Start from the lower bound of the interval — €250 of discounted margin per new subscriber — and not from the constant rate's €103, which, if the model is correct, would lead to rejecting profitable campaigns. A ceiling equal to the expected value corresponds to break-even, not to a profit: the safety margin, notably for model error and the choice of discount rate, remains a decision for the finance department. This ceiling also assumes that the next recruits behave like this cohort [C37]; to be checked on each new cohort and each new channel.
  • Retention. The value of the remaining subscribers — about €494,000, between €420,000 and €568,000 across the 1,058 subscribers — measures what is at stake, not a budget. The retention budget depends on the gain that a retention action actually produces: to be measured with an experiment. The sensitivity of value to retention, its elasticity, is calculated with the same model [C36], and an aggregate rate underestimates it too [C27].
  • Monitoring. Re-estimate the model every month, as the cohort ages, and compare each new projection with what is observed.

The decision remains a human one and commits the finance department.

15

When to use it, when not to

  • Use: subscriptions with fixed renewal points, monthly or annual, whenever retention must be projected beyond the history: value of a new subscriber, value of the base, acquisition ceiling.
  • Do not use for targeting: to find out which subscriber is likely to leave next month, churn scoring models, which use variables such as the number of calls to customer service, are the right tool; they are, however, poorly suited to projecting the survival curve, since the future values of these variables are unknown [C14]. See the page on predicting churn with logistic regression.
  • Do not use alone: with an introductory offer or strong seasonality [C29], or if individual loyalty is suspected to change with tenure [C17, C28].
  • Alternatives: in continuous time, the exponential-gamma model [C15]; without subscriptions, the Pareto/NBD or BG/NBD models [C16] — see the page on estimating noncontractual CLV; to describe observed survival without extending it, the Kaplan-Meier estimator — see the page on analyzing retention and survival. To judge a projection on held-out months, see the page on validating a marketing forecast.

16

Reproducible implementations

  • Python 3.14.5, standard library only: msc-p043-reference.py, run with python msc-p043-reference.py
  • R 4.6.1, base only: msc-p043-reference.R, run with Rscript msc-p043-reference.R
  • Both programs print exactly the same lines, whose sha256 fingerprint, after line-ending normalization, is f98cbaeade192173a375ae7e4290cb24ec7561c498f7fbca17f45811847186fb. Declared seed 20261002. No third-party library: the software does not validate the method, it reproduces it.
  • External check: before any calculation on the synthetic data, both programs recover the estimates published by Fader and Hardie and the residual lifetimes in their 2010 table, and stop with an error otherwise [C12, C13, C32].

17

Deliverable

The projected retention curve and its out-of-sample validation, the parameters α and β with their standard errors, the value of a new subscriber and that of the remaining subscribers with their intervals, the share of these values lying beyond the verified horizon, their sensitivity to the discount rate, the comparison with the constant rate, and the list of assumptions to check before deciding.

18

References and level of evidence

Method connections

Parent territoryCustomer and choice science: behavior, value and heterogeneityCompare withHow do you estimate CLV with BG/NBD and Gamma-Gamma?Compare withHow do you analyze retention with a survival model?Compare withWhich customers have the highest probability of churn?RequiresHow should uncertainty in a marketing result be expressed?ValidatesHow do you validate a marketing forecast?

Read next

MSC-H-004Customer and choice science: behavior, value and heterogeneity→MSC-P-028How do you estimate CLV with BG/NBD and Gamma-Gamma?→MSC-P-030How do you analyze retention with a survival model?→MSC-P-029Which customers have the highest probability of churn?→MSC-P-003How should uncertainty in a marketing result be expressed?→MSC-P-033How do you validate a marketing forecast?→