Research & Evidence

Search for a method

Search titles, questions, territories and MSC identifiers.

    ← All methods
    METHOD DOSSIERMSC-P-046Customer scienceVerified scientific dossier

    How much will a customer spend per purchase? The Gamma-Gamma model

    Gamma-Gamma model: predict each customer's average spend per purchase from past purchases, checked on real CDNOW data and on a held-out period.

    Scientific editorial team: Marketing Science Center

    Direct answer

    Predict each customer's average spend per purchase from the number and value of their purchases, with an interval, and validate the prediction on a held-out period against simple averages.

    Gamma-Gamma model: predict each customer's average spend per purchase from past purchases, checked on real CDNOW data and on a held-out period.

    Fader, Hardie & Lee, 2005Fader & Hardie, 2013

    01

    In short

    A customer's value rests on two things: how many times they will buy, and how much they will spend each time. The Gamma-Gamma model of Fader, Hardie and Lee deals with the second [C02]: it forecasts each customer's average spend per purchase, the average order value, from their past purchases [C01]. Its rule: the forecast pulls a customer's observed mean toward that of all customers, all the more strongly as the customer has bought little [C20, C21].

    The page's program applies it to the real data published by the authors: 2,357 customers of the online music retailer CDNOW, acquired in the first quarter of 1997 [C03, C04]. It recovers the results the authors published on these data, except one descriptive statistic. It then judges the forecasts on thirty-nine weeks the estimation has not seen, for the 491 customers who buy again:

    • the authors' model, which uses repeat purchases only, is off on average by $16.37 per purchase; that is clearly better than the mean of those repeat purchases alone, $17.35, but no better than the mean of all the customer's purchases, the first included, $16.09;
    • the same model estimated on all purchases, a variant added by the page after a first look at these weeks, is off by $15.78: slightly better than that simple mean, but with no clear difference from the authors' model. It remains to be confirmed on other data.

    The model also underestimates how much a customer's purchases vary around their mean, and it corrects small baskets too little. It serves to correct means based on few purchases; it does not forecast the total for the base any better.

    02

    The situation

    An online music retailer acquired its customers in the first quarter of 1997. By 30 September 1997, 946 of the sample's 2,357 customers had made at least one repeat purchase [C07]. To put a figure on their value, the marketing team already has a model that forecasts how many purchases each will make, such as the one on the page on estimating noncontractual CLV; what it lacks is the amount of those purchases.

    Two customers puzzle the team. The first made a single repeat purchase, of $100; the second made six, at $20 on average. Should it forecast $100 for the first, $20 for the second, or something else?

    The page's data are CDNOW's, as Bruce Hardie publishes them: a one-in-ten systematic sample of the cohort [C03], that is, 6,919 rows giving the date of purchase, the number of CDs and the amount in dollars [C05].

    03

    The scientific question

    What average order value should we expect from a customer, given the number x of their repeat purchases and their average amount z̄? The observed mean is only an imperfect estimate of the customer's true mean, which cannot be seen [C08]; it must therefore be corrected, we must say by how much, and we must check that the correction forecasts better than the simple means.

    04

    Why the simple method fails

    Three common shortcuts:

    • Taking the customer's past mean. The authors ask whether a model is needed and answer that this mean cannot necessarily be trusted [C09]: if the average spend of all customers is $35, what should we forecast for a customer whose only repeat purchase cost $100 [C10]? Over the held-out weeks, customers whose mean repeat purchase was below $20, $14.84 on average, then spent $26.82 per purchase; those whose mean reached or exceeded $50, $79.07 on average, spent $69.20. The extreme values of a small number of purchases tend not to repeat. Counting the first purchase in the mean already softens this flaw.
    • Taking the mean of all customers. It erases real differences between customers: over the held-out weeks, it is the worst of the forecasts compared.
    • Assuming a normal distribution, as Schmittlein and Peterson do: it allows negative spending and it is symmetric [C11], whereas spending is skewed, with a long tail toward large amounts [C12]. Here, the mean order value, $35.08, exceeds the median, $27.50, which exceeds the mode, $14.96.

    Multiplying the forecast number of purchases by the customer's observed mean is tempting; the authors write that it ignores regression to the mean [C46, C47].

    05

    The intuition

    The model rests on three substantive assumptions [C14, C15, C16]: the amount of each of a customer's purchases varies randomly around that customer's own mean; these means differ from one customer to another but do not change over time; and their distribution across customers does not depend on purchase frequency. It adds two choices of functional form, detailed below.

    From the whole base, the model learns two things: how much a given customer's purchases vary around their mean, and how much the means vary from one customer to another. For each customer, the forecast is a weighted average of the population mean and the customer's observed mean [C20]; the more the customer has bought, the more their own mean weighs [C21]. On these data, a single repeat purchase already gives the customer's mean a weight of 0.69; four repeat purchases, 0.90.

    06

    The data you need

    For each customer, over the estimation period:

    • the number x of their repeat purchases, that is, not counting the first;
    • the average amount z̄ of these repeat purchases: the authors' descriptive table covers the average value of repeat purchases [C07, C24].

    Purchases made on the same day are added together and count as a single purchase; this convention exactly recovers the authors' 946 customers and descriptive table. Customers without a repeat purchase do not enter the authors' estimation; their forecast is the population mean [C48]. The page's variant also counts the first purchase: all customers enter it, except 8 whose single purchase is recorded at $0 in the file.

    You then need to check the independence assumption: the authors measure a correlation of 0.11 between average spend and the number of purchases [C38], which they do not consider strong enough to call the model into question [C41].

    07

    The formal model

    This section is for practitioners. Decision-makers can skip it and resume at "The declared calculation": the "In plain words" sentence at its end is enough.

    • z₁, …, z_x: the amounts of a customer's x repeat purchases; z̄ their mean. Z denotes the amount of a purchase; a hat, as in p̂, denotes an estimated value.
    • Each amount follows a gamma distribution with shape p and rate ν (the authors say "scale"; the mean is ζ = p/ν, the customer's true average spend) [C17]. The parameter p is the same for all: all customers have the same coefficient of variation, the standard deviation of a purchase relative to the customer's mean, 1/√p [C19].
    • ν varies across customers according to a gamma distribution with shape q and rate γ [C18]; these two distributional assumptions make up the Gamma-Gamma model [C53]. The true mean ζ = p/ν then follows an inverse gamma distribution. The gamma distribution has properties close to those of the lognormal, with a slightly thinner tail; unlike the lognormal, it yields closed-form expressions [C13]. The model adapts that of Colombo and Jiang [C54].
    • Distribution of the observed mean: f(z̄ | x) = Γ(px + q) / (Γ(px) Γ(q)) × γ^q z̄^(px−1) x^(px) / (γ + x z̄)^(px+q), where Γ is the gamma function, which extends the factorial, not to be confused with the parameter γ. This formula gives the probability density of a mean z̄ observed after x purchases.
    • Population mean: E(Z) = pγ / (q − 1).
    • Forecast for a customer: E(Z | z̄, x) = p(γ + x z̄) / (px + q − 1), that is, the weighted average (1 − w) E(Z) + w z̄, with a weight w = px / (px + q − 1) given to the customer's mean [C20, C21].
    • Estimation: we keep the values of p, q and γ that make the observed means of the 946 customers most likely, that is, the maximum of the log-likelihood, the sum of the ln f(z̄ | x), found by a standard numerical method [C22].

    In plain words: each customer has a true average spend, which their first purchases only partly reveal; as long as they have bought little, we forecast between their mean and everyone's, and closer to their own with each purchase.

    08

    The declared calculation

    1. Data. The published file, read as is: 6,919 rows, 2,357 customers, $244,091.94; with purchases made on the same day added together, 6,696 purchase days remain. Estimation on weeks 1 to 39, up to 30 September 1997; weeks 40 to 78, from 1 October 1997 to 30 June 1998, held out, as in the authors' work [C06].
    2. Check against the published values. Before any result, both programs recover the 946 customers who made a repeat purchase [C07], the eight rows of the note's descriptive table [C24, C25, C26, C27, C28, C29, C30, C31], the three parameters [C23], the nine-cent gap between theoretical and observed mean [C34], the two modes [C35, C36], the correlations [C38, C39, C40], the article's log-likelihoods [C42] and the mean over 78 weeks [C44]. Otherwise they stop with an error. Quartiles are computed by Weibull's rule, the position (n + 1) × probability, the only common rule that recovers both published values.
    3. Estimation of the maximum likelihood on the logarithms of p, q and γ, by the Nelder-Mead simplex, an algorithm that searches for the best fit step by step, restarted until stable; the programs then check that the slope of the likelihood is zero there. The log-gamma function is written out in full in both languages.
    4. The page's variant: the same model, estimated on all purchases in weeks 1 to 39, the first included. It was added after a first look at the held-out weeks, following a review: its validation is therefore not blind, and its standard errors do not account for this choice.
    5. Validation on the held-out weeks: five forecasts of average spend per purchase are compared with what the customers spent: the mean of repeat purchases, the mean of all purchases, the population mean, the authors' model and the variant. The error is observed spend minus the forecast; when positive, it signals an underforecast. Each difference is given with its standard error, which measures the imprecision due solely to the choice of customers; a difference of more than two standard errors is judged clear. Differences are computed before rounding.
    6. A customer's dispersion around their mean: for customers with at least three repeat purchases, the observed coefficient of variation of their purchases is compared with the one given by 200 simulations of the model, with the same numbers of purchases.
    7. Uncertainty: bootstrap, which redraws at random, with replacement, 1,000 times 946 customers out of the 946, and re-estimates the model each time; interval from the 25th to the 975th sorted value. These intervals are described as nominal 95%: that is the level the method aims at, which nothing here guarantees exactly. The pseudo-random number generator is written out in full (seed 20261007; s ← 1103515245 · s + 12345 mod 2³¹), so that Python and R draw the same numbers.
    8. Invariants: 19 identities, including the distribution of z̄ integrating to 1, its mean equal to pγ / (q − 1) and the forecast equal to the weighted average; the programs stop if one fails.

    09

    The complete worked example

    The published results, recomputed. Average value of repeat purchases per customer, weeks 1 to 39, in dollars:

    StatisticRecomputedPublished
    Minimum2.992.99 [C24]
    First quartile15.74515.75 [C25]
    Median27.497527.50 [C26]
    Third quartile41.795641.80 [C27]
    Maximum299.6338299.63 [C28]
    Mean35.077835.08 [C29]
    Standard deviation30.283530.28 [C30]
    Mode14.9614.96 [C31]

    The parameters are recovered: p̂ = 6.2496, q̂ = 3.7442, γ̂ = 15.4435, against 6.25, 3.74 and 15.44 published [C23]. Excel's Solver tool, in the authors' workbook, stops slightly elsewhere at the fourth decimal, because the likelihood surface is nearly flat; the maximum reached here is very slightly higher. The theoretical mean, $35.1704, exceeds the observed mean by $0.0925, the authors' nine cents [C34]. The mode of the fitted distribution falls at $19, that of the data at $14.96, as published [C35, C36]. The correlation between average spend and number of purchases is 0.1139, and 0.0570 without the customer with 21 purchases and a $299.63 mean; the test that this correlation is zero then gives a p-value of 0.0801. These are the published 0.11, 0.06 and 0.08 [C38, C39, C40].

    Two gaps remain.

    • The log-likelihood. At the maximum, the note's formula gives −4,055.92. The article publishes −4,659 [C42]. The gap is exactly the sum of the ln x of the 946 customers, 603.37: −4,055.92 − 603.37 = −4,659.29. The article therefore probably computed the likelihood of total spend x z̄, although the formula it gives is that of the mean z̄; the maximum, and hence the parameters, are the same. With this convention, the parameters estimated over the 78 weeks give −4,661.41 on weeks 1 to 39, the published −4,661 [C42], and a population mean of $35.81, the authors' "$36" [C44].
    • Skewness. The article gives a skewness of 4 and a kurtosis of 17 [C32, C33]. The page finds 17.15 for kurtosis, as excess over the normal distribution, but 3.40 for skewness; no variant of the definition we tried gives 4.

    The fitted model. The customers' true means have a mean of $35.17 and a standard deviation of $26.63. The model estimates the coefficient of variation of a customer's purchases around their own mean at 0.40; what follows shows that the data contradict it.

    The forecast for a customer, in dollars per purchase, according to the authors' model:

    Repeat purchasesWeight of the customer's meanObserved mean $20$50$100
    10.694924.6345.4880.22
    20.820022.7347.3388.33
    40.901121.5048.5393.59
    80.948020.7949.2396.63

    For the customer with a single $100 repeat purchase, we expect $80.22 per purchase; for the one with six repeat purchases at $20, $21.03. The weight of the customer's mean reaches 90%, a threshold chosen by the page, at the fourth repeat purchase; with the parameters estimated over 78 weeks, at the seventh. The authors judge, without a numerical threshold, that seven or eight purchases are needed before trusting the observed mean [C45].

    The page's variant, estimated on all purchases in weeks 1 to 39: p̂ = 9.8813, q̂ = 3.3988, γ̂ = 7.9286, population mean $32.66. For a customer whose only purchase cost $100, the variant expects $86.85, with a weight of 0.80 given to their mean. Within the same customer, the first purchase is about equal to the repeat purchases: $36.14 on average for the 946 customers who bought again, against $35.08 for the mean of their repeat purchases, with a median ratio of 0.97. By contrast, the 1,411 customers who did not buy again had made a first purchase of only $30.88: customers whose first purchase was smaller bought again less. The variant mixes these two groups, which lowers its population mean.

    Validation on weeks 40 to 78. Of the 946 customers who made a repeat purchase, 491 buy again. Error on their average spend per purchase, in dollars, and difference in absolute error from the mean of all purchases:

    ForecastMean absolute errorRoot mean squared errorMean errorDifference from the mean of all purchases
    Mean of repeat purchases17.3529.652.30+1.26 (standard error 0.45)
    Mean of all purchases16.0928.282.05—
    Population mean18.9431.622.67+2.85 (0.91)
    Authors' model16.3728.062.17+0.28 (0.41)
    The page's variant15.7827.832.19−0.31 (0.09)

    The authors' model reduces the error by 5.6% compared with the mean of the customer's repeat purchases alone, that is, $0.98 per purchase (standard error 0.25), but does no better than the mean of all purchases: its difference, +$0.28, is smaller than its standard error. The variant does better than this simple mean, by 1.9%, that is, $0.31 per purchase; compared with the authors' model, its advantage, $0.59 (standard error 0.37), is not clear. It is an average gain: the variant beats the simple mean for 285 customers out of 491, that is, 58.0%.

    For 193 other customers, who made no repeat purchase during estimation but buy afterwards, the value of their single purchase is off by $20.25, the population mean, which is the authors' model's forecast, by $20.92, the variant by $18.75. The variant's advantage is clear against the single purchase, $1.51 (standard error 0.44), but not against the population mean, $2.17 (standard error 1.72). Across the 684 customers who buy during these weeks, the error is $17.27 for the mean of all purchases, $17.65 for the authors' model and $16.62 for the variant; the difference between the variant and the authors' model, $1.03 (standard error 0.55), is not clear.

    Regression to the mean, by band of mean repeat purchase, forecasts of the authors' model and of the variant:

    Band of mean repeat purchaseCustomersMean repeat purchaseAuthors' modelVariantSpend observed afterwardsObserved minus authors' model
    Under $2014914.8419.4019.6526.827.42 (standard error 1.61)
    $20 to $5025532.7833.1632.9133.570.41 (1.10)
    $50 and over8779.0770.8471.0369.20−1.64 (5.70)

    Both models pull extreme means in the right direction, but not far enough: small baskets are clearly underforecast, by $7.42 per purchase; large baskets are overforecast by $1.64, within the margin of chance.

    The total for the base. Multiplied by the number of purchases actually made, the authors' model's forecast gives $68,181.97 of spending over weeks 40 to 78, against $70,976.39 observed, that is, −3.9%; the variant, −4.6%; the mean of all purchases, −4.2%. By number of repeat purchases during estimation, average total spend per customer over these weeks:

    Repeat purchasesCustomersObservedForecast by the authors' modelDifference
    01,4118.308.33−0.03 (standard error 0.55)
    143927.4223.164.26 (1.82)
    221451.9253.72−1.80 (3.95)
    310063.1956.236.96 (4.47)
    46290.98100.75−9.77 (6.54)
    538124.45105.7018.75 (13.68)
    629127.17126.740.43 (8.29)
    7 or more64246.04237.638.42 (12.84)

    Only the single-repeat-purchase row departs beyond chance: these customers are underforecast by $4.26. The authors, who give no standard error, judge by eye that no particular bias calls their assumptions into question [C50].

    A customer's dispersion around their mean. For the 293 customers with at least three repeat purchases, the median coefficient of variation of their purchases is 0.47. Computed on a few purchases, a coefficient of variation is on average smaller than the true one: the model therefore expects 0.36, not 0.40, and in 95% of the 200 simulations, between 0.34 and 0.38. A customer's purchases therefore vary clearly more than the model assumes, especially when their spend is high: 0.38 for the 61 customers under $20, 0.48 for the 180 from $20 to $50, 0.50 for the 52 at $50 and over.

    10

    Validity assumptions

    • Spend independent of frequency [C16]. Here, the correlation is 0.11, halved when a single customer is removed [C38, C39], and the authors do not see it as a serious violation [C41]. But customers who do not buy again had made a smaller first purchase, $30.88 against $36.14 (a gap of $5.26, standard error 1.47): a sign that spend and frequency are not entirely independent. For products that customers stock up on, one may expect an even stronger link between the pace of purchases and their amount [C51]. Do your big spenders also buy more often, or less often?
    • An average spend that is stable over time for each customer [C15]. The log-likelihood barely changes when moving from 39 to 78 weeks [C42, C43], but the parameters shift: the weight of a single repeat purchase falls from 0.69 to 0.59, and the customer with a single $100 repeat purchase goes from $80.22 to $73.99. The average level, however, has changed little: a repeat purchase was worth $38.81 on average during estimation, a purchase $37.71 afterwards. Have your prices, your offer or your customers changed since the estimation period?
    • The same coefficient of variation for all customers [C19]. The data contradict it: a customer's purchases vary more than forecast, especially among big customers. This is one lead to explain the insufficient regression to the mean, but it does not explain the underforecast of small baskets, whose dispersion is close to the model's. Are your big customers' baskets more irregular than small customers'?
    • A gamma distribution of amounts, without price points. The model knows nothing of round or psychological prices ending in .99: it places the mode at $19 when the data place it at the price of one CD, $15 [C35, C36, C37]. The authors preferred a simple model to a better fit [C52]. Do your baskets cluster on a few prices?
    • Repeat purchases only, in the authors' model: the first purchase does not enter the mean, and customers without a repeat purchase receive the population mean [C48]. What share of your customers has not yet bought again?
    • Testable only indirectly: the inverse gamma distribution of the true means, since a true mean is never observed, only means of a few purchases; it is judged by the fit of the distribution of observed means. Not testable here: that a customer's mean stays stable beyond 30 June 1998.

    11

    Diagnostics and uncertainty

    • Fit: compare the distribution of observed means with the one the model forecasts. Here, the theoretical mean departs from the observed one by only $0.0925, but the forecast mode, $19, exceeds the observed one, $14.96 [C35].
    • Independence: compute the correlation between average spend and number of purchases, with and without extreme customers [C38, C40], and compare the first purchase of those who buy again with that of those who do not.
    • A customer's dispersion: compare the observed coefficient of variation of each customer's purchases with the model's; here 0.47 against 0.36.
    • Out-of-period validation: compare the model with the simple means, including the mean of all purchases, with the standard error of each difference. Weighted by the number of purchases made afterwards, the mean absolute error keeps the same ranking: $13.44 for the variant, $13.60 for the authors' model, $13.80 for the mean of all purchases, $14.61 for the mean of repeat purchases.
    • Estimation uncertainty, by bootstrap over the 946 customers, nominal 95% intervals: the population mean ranges from $33.26 to $37.02; the forecast for the customer with a single $100 repeat purchase, from $72.95 to $88.58; the weight of a single repeat purchase, from 0.59 to 0.82; the coefficient of variation, from 0.30 to 0.46; the standard deviation of the true means, from $21.75 to $32.58. The data separate p poorly from γ: p ranges from 4.68 to 10.85 and γ from 7.42 to 24.47, whereas the forecasts, which combine them, remain more stable. These intervals concern average forecasts, not what a customer will actually spend: a single purchase varies far more.
    • What the validation does not check: it covers only customers who buy again, each weighing the same whatever their number of purchases; it uses the number of purchases actually made over weeks 40 to 78 and does not judge the forecast of that number, which is the job of another model [C49]. These weeks include the 1997 year-end holidays, which could raise spending, something the average level does not show. Nothing is said beyond 30 June 1998.
    • What these data cannot settle: one retailer, one product, customers acquired in 1997; the result does not carry over without checking.

    12

    Interpretation

    Three lessons. First, a customer's mean repeat purchase is a poor forecast when they have bought little: a single $100 repeat purchase points to about $80, not $100. The model quantifies this regression to the mean instead of guessing it.

    Second, on these data, the authors' model adds nothing compared with a simple mean of all purchases, the first included: it beats the mean of repeat purchases alone because it corrects, but it does without the information in the first purchase. Fed all purchases, the same model does slightly better than this simple mean, on average across customers; but it was chosen after the fact, and its advantage over the authors' model is not clear. None forecasts the total for the base any better: the gain lies in the value of each customer, not in overall revenue.

    Finally, the model corrects too little: customers under $20 then spent $26.82 per purchase, when it forecast 19.40. These data do not reveal the cause. A per-customer dispersion greater than expected is measured here, but it mainly concerns big customers; price points, the shape of the distribution of true means, the link between spend and frequency or a change over time all remain possible.

    13

    Allowed and forbidden conclusions

    • Allowed: "For a CDNOW customer who made a single repeat purchase, of $100, the authors' model forecasts an average spend of about $80 per purchase for their next purchases, from $72.95 to $88.58 at a nominal 95%, counting estimation uncertainty only."
    • Allowed: "Over weeks 40 to 78, the model estimated on all purchases, a variant chosen after a first look at these weeks, is off on average by $15.78 per purchase, against $16.09 for the mean of all the customer's purchases; this remains to be confirmed on other data."
    • Forbidden: presenting the model's forecast as the value of a customer. It forecasts only the amount per purchase; it must be combined with a model of the number of purchases, as the authors do [C49], then with a margin and a discount rate [C55, C56].
    • Forbidden: claiming that the authors' model forecasts better than the customer's past mean without saying which mean: it beats the mean of repeat purchases, not the mean of all purchases.
    • Forbidden: concluding that a marketing action would make customers spend more. The model describes spending; it does not measure the effect of any action.
    • Forbidden: applying CDNOW's parameters to another firm, or applying the model without having checked that your customers' spending does not depend on their purchase frequency [C51].

    14

    Possible marketing decision

    • Value each customer with a corrected forecast, not with the mean of their repeat purchases alone; failing a model, the mean of all their purchases does almost as well. The model estimated on all purchases can be considered, then validated on your own data, over a period you have not yet looked at. Then multiply by the expected number of future purchases, each brought back to its value today, and by the margin, as the authors do, who use a 30% margin and a 15% annual discount rate [C49, C55, C56].
    • Neither overestimate a single large purchase nor underestimate a small one: a single $100 repeat purchase is worth a forecast of $80.22 according to the authors' model, and a customer with a $20 mean after a single repeat purchase, $24.63; the data even show that the correction should be stronger for small baskets.
    • Check on your own data, before using it, the independence between spend and frequency, the dispersion of each customer's purchases and the validation against the mean of all purchases; re-estimate after any price change.

    The decision remains human and involves the marketing department and, for customer value, the finance department.

    15

    When to use it, when not to

    • Use it: for repeat purchases of varying amounts, outside subscriptions, when you want the value of each customer; alongside a model of the number of purchases, such as the one on the page on estimating noncontractual CLV or the one on the page on the value of a seasonal customer.
    • Use it with caution when spend depends on frequency, for example for products that people stock up on [C51]: you then need a model that links the number of purchases and their amount instead of assuming them independent. Caution also when prices cluster on a few values [C37], or when a customer's purchases are very irregular.
    • Do not use it for a fixed-price subscription, where the amount is known: see the page on subscriber value; nor to value the whole base including future customers: see the page on the value of the customer base.
    • Do not use it to measure the effect of an action on spending: that requires an experiment.
    • To express the uncertainty of a result, see the page on expressing uncertainty; to judge a forecast on a held-out period, the one on validating a marketing forecast.

    16

    Reproducible implementations

    • Python 3.14.5, standard library only: msc-p046-reference.py, run with python msc-p046-reference.py
    • R 4.6.1, base only: msc-p046-reference.R, run with Rscript msc-p046-reference.R
    • The data are not redistributed: download the CDNOW sample from brucehardie.com/datasets and place CDNOW_sample.txt next to the programs.
    • Both programs print exactly the same lines; their sha256 fingerprint, after line-ending normalisation, is recorded in the dossier manifest. Declared seed 20261007. No third-party library: the software does not validate the method, it reproduces it.

    17

    Deliverable

    For each customer, the number of purchases, their average amount and the corrected forecast with its weight; the estimated parameters and their interval; the check of independence between spend and frequency and that of each customer's dispersion; validation on a held-out period, against the mean of all purchases, with standard errors; and the assumptions to check before deciding.

    18

    References and level of evidence

    • Fader, P. S., Hardie, B. G. S. & Lee, K. L. (2005). RFM and CLV: Using Iso-Value Curves for Customer Base Analysis. Journal of Marketing Research 42(4), 415–430. — the authors' version posted by Bruce Hardie, full text verified; Schmittlein and Peterson (1994) and Colombo and Jiang (1999) are cited in it.
    • Fader, P. S. & Hardie, B. G. S. (2013). The Gamma-Gamma Model of Monetary Value. Technical note. brucehardie.com/notes/025 — full text verified.
    • CDNOW sample and its documentation, published by Bruce Hardie: brucehardie.com/datasets (CDNOW_sample.zip); the full cohort numbers 23,570 customers [C04].
    • Level of evidence: real data, one retailer, 1997 and 1998; published results recovered from the data, except skewness; validation on 39 held-out weeks, with a variant added after a first look at those weeks; reproducible in Python and R. This is not proof that the model suits your firm.

    Method connections

    Parent territoryCustomer and choice science: behavior, value and heterogeneityCompare withHow do you estimate CLV with BG/NBD and Gamma-Gamma?Compare withWhat is a customer who buys at most once a season worth?Compare withWhat are a firm’s current and future customers worth?RequiresHow should uncertainty in a marketing result be expressed?ValidatesHow do you validate a marketing forecast?

    Read next

    MSC-H-004Customer and choice science: behavior, value and heterogeneity→MSC-P-028How do you estimate CLV with BG/NBD and Gamma-Gamma?→MSC-P-044What is a customer who buys at most once a season worth?→MSC-P-045What are a firm’s current and future customers worth?→