Research & Evidence

Search for a method

Search titles, questions, territories and MSC identifiers.

40 results
  1. MSC-P-001How do you turn a marketing claim into a testable question?Decision Science
  2. MSC-P-002Correlation or causality: what can an analysis actually support?Marketing Measurement
  3. MSC-P-003How should uncertainty in a marketing result be expressed?Decision Science
  4. MSC-P-004Statistical significance or effect size: which result should be interpreted?Decision Science
  5. MSC-P-005How do you measure a marketing construct that is not directly observable?Market Research
  6. MSC-P-006How do you design and validate a measurement scale?Market Research
  7. MSC-P-007Alpha or omega: how should scale reliability be assessed?Market Research
  8. MSC-P-009PCA, EFA or CFA: which method should you choose?Market Research
  9. MSC-P-010When should you run a marketing experiment?Marketing Measurement
  10. MSC-P-011How do you design an A/B test that actually estimates an effect?Marketing Measurement
  11. MSC-P-012How many observations does an experiment need?Decision Science
  12. MSC-P-013How do you measure campaign incrementality with a control group?Marketing Measurement
  13. MSC-P-017How do you detect selection, contamination and attrition in an experiment?Marketing Measurement
  14. MSC-P-018Predictive or causal regression: what are you trying to estimate?Marketing Models
  15. MSC-P-019How do you diagnose a marketing regression before interpreting it?Marketing Models
  16. MSC-P-022How do you estimate price elasticity and its uncertainty?Pricing Science
  17. MSC-P-026Logit vs Probit: how do you choose for purchase probability?Customer Science
  18. MSC-P-029Which customers have the highest probability of churn?Customer Science
  19. MSC-P-027TAM, UTAUT or UTAUT2: which framework should be used to study technology acceptance?Market Research
  20. MSC-H-001Measurement and causality: how can a marketing effect be established?Marketing Measurement
  21. MSC-H-002Marketing response models: shape, delay and saturationMarketing Models
  22. MSC-H-003Pricing science: connecting price, demand and contributionPricing Science
  23. MSC-H-004Customer and choice science: behavior, value and heterogeneityCustomer Science
  24. MSC-H-005Measurement science: building valid indicatorsMarket Research
  25. MSC-H-006Statistical decision methods: choose, quantify, validateDecision Science
  26. MSC-P-008How do you validate a marketing measurement scale?Market Research
  27. MSC-P-014How do you design a marketing geo experiment?Marketing Measurement
  28. MSC-P-015How do you estimate an effect with difference-in-differences?Marketing Measurement
  29. MSC-P-020How do you address price endogeneity?Pricing Science
  30. MSC-P-021Fixed or random effects: which panel model should you choose?Marketing Models
  31. MSC-P-023How do you estimate a demand function?Pricing Science
  32. MSC-P-024How do you simulate a price-volume-margin scenario?Pricing Science
  33. MSC-P-028How do you estimate CLV with BG/NBD and Gamma-Gamma?Customer Science
  34. MSC-P-030How do you analyze retention with a survival model?Customer Science
  35. MSC-P-031How do you build a useful customer segmentation?Customer Science
  36. MSC-P-032How do you test segmentation stability?Customer Science
  37. MSC-P-033How do you validate a marketing forecast?Decision Science
  38. MSC-P-034How do you build a Monte Carlo simulation for a marketing decision?Decision Science
  39. MSC-P-035How do you model saturation and adstock?Marketing Models
  40. MSC-P-039Which statistical test should you choose?Decision Science
All methods
METHOD DOSSIERMSC-P-021Regression and econometrics

Fixed or random effects: which panel model should you choose?

Fixed effects, standard random effects and CRE/Mundlak do not answer the same question. Choice depends on the within or between estimand, orthogonality and error structure, not an automatic test.

Direct answer

Separate the within-entity association from the between-entity association and choose a decision-aligned specification.

Fixed effects, standard random effects and CRE/Mundlak do not answer the same question. Choice depends on the within or between estimand, orthogonality and error structure, not an automatic test.

Hausman, 1978Bell, Fairbrother & Jones, 2019

01

The answer in 30 seconds

1

Fixed effects estimate a within-unit association from changes observed inside each entity.

2

Standard random effects combine within- and between-unit variation under a stronger orthogonality assumption.

3

The CRE/Mundlak model separates both contrasts. It is often more informative than an automatic Hausman-based choice.

Bell, Fairbrother & Jones (2019) Wooldridge (2019)

02

Three reading levels

  1. 1

    Decision-maker: ask whether the decision concerns change within the same store or persistent differences across stores.

  2. 2

    Practitioner: check identifier, time, within variation, missingness and the clustering level for standard errors.

  3. 3

    Analyst: declare the estimand, state FE, RE and CRE assumptions, then compare coefficients, uncertainty and residuals.

03

Concrete marketing situation

A network follows 12 stores for 8 weeks. For each store-week it observes the share of visitors exposed to display and orders per 1,000 visits. Some stores persistently have more display and fewer orders. The operational question, however, concerns what accompanies a display increase within the same store.

04

Scientific question and estimand

Unit: store × week. Primary estimand: within coefficient βW, the average change in orders per 1,000 visits associated with a +1-point display share within the same store. Secondary estimand: between coefficient βB across store means. These are panel associations, not identified causal effects.

βW = Cov(xit − x̄i, yit − ȳi) / Var(xit − x̄i)

05

Why the simple model can mislead

  • A pooled regression can conflate change within a store with structural differences across stores.
  • Standard RE becomes a hard-to-interpret blend when within and between associations differ.
  • Clustered standard errors address residual dependence, not identification bias.

Bell, Fairbrother & Jones (2019) Hausman (1978)

06

Intuition: separate within and between

Decomposing xit into store mean x̄i and deviation xit − x̄i makes both questions visible. The within term compares a store with itself. The between term compares persistently different stores. Standard RE imposes equality; CRE/Mundlak estimates them separately.

xit = (xit − x̄i) + x̄i

Wooldridge (2019)

07

Required data

  • One row per entity and period, with a stable identifier, date or period, outcome and time-varying predictors.
  • Enough within-entity variation and enough entities for defensible clustered uncertainty.
  • Timing, attrition, definition changes and missing values documented before estimation.

08

Three models, three assumptions

Fixed effects

yit − ȳi = βW(xit − x̄i) + (εit − ε̄i)

The example is one-way store FE. It absorbs stable store heterogeneity, not common week shocks.

Standard RE

yit = α + βRE xit + ui + εit

Assumes E[ui|Xi]=0 and one common within and between coefficient.

CRE / Mundlak

yit = α + βW(xit−x̄i) + βB x̄i + ai + εit

Separates both contrasts and permits a Wald test of βB−βW=0 with a declared covariance. The example does not run that test.

i indexes store, t week, y orders per 1,000 visits, x display share, x̄i and ȳi store means, α the intercept, ui or ai stable heterogeneity, εit the idiosyncratic shock, βW the within contrast and βB the between contrast. Two-way FE adds λt, a week effect; βW then uses variation net of common week shocks.

Two-way FE: yit = αi + λt + βW xit + εit

09

Declared calculation, step by step

  1. 01

    Compute x̄i and ȳi for each store, then within deviations.

  2. 02

    Estimate βW on demeaned variables and cluster uncertainty by store.

  3. 03

    Estimate βB on store means and standard RE by GLS quasi-demeaning.

  4. 04

    Estimate CRE with within deviation and mean; compare βW, βB, βRE and their uncertainty.

10

End-to-end numerical example

MSC-P-021-PANEL
Synthetic dataset created for learning. It describes no real store network.

Reference results on 96 rows, 12 stores × 8 weeks
SpecificationCoefficientMeaning
Pooled−1.140Within/between blend
Fixed effects−0.600Within association
Between−1.200Association across means
Standard RE−0.689Declared Python moments estimator
CRE / MundlakβW=−0.600; βB=−1.200Separated contrasts
Download the synthetic CSV

βW

−0.600

Clustered SE

0.023

95% t interval, df=11

[−0.651; −0.549]

Declared teaching convention: store-cluster sandwich, finite G/(G−1) correction, t(11)=2.200985 critical value. With only 12 clusters, this interval illustrates the declared calculation but does not guarantee reliable small-sample inference. It covers no identification bias.

11

Validity assumptions

  • FE: strict exogeneity of the idiosyncratic error conditional on predictor history, plus sufficient within variation.
  • Standard RE: the same condition, plus orthogonality between the stable entity effect and included predictors.
  • CRE/Mundlak: the entity mean adequately represents correlation between stable heterogeneity and predictors under the declared specification.
  • None of these models automatically removes time-varying confounding, reverse causality, measurement error or anticipation.

Operational questions: does display respond to past, current or anticipated order shocks? Do price, stock or traffic change simultaneously? Are stores without variation excluded? Is the store mean sufficient to represent stable correlation?

Wooldridge (2019) Bell, Fairbrother & Jones (2019)

12

Diagnostics and uncertainty

01

Observed structure: 96/96 complete rows, 12 stores × 8 weeks, balanced panel, no store without variation. Within-store display deviations range from −3.5 to +3.5 points (population SD 2.291). The FE coefficient is therefore identified by within-store variation present in every store.

02

Uncertainty: only 12 clusters. The store-sandwich SE is 0.022972 with G/(G−1) correction; the t(11) interval is [−0.650562, −0.549438]. It illustrates this teaching convention without guaranteeing reliable inference with few clusters. The store-only fixed-effects model does not absorb common week shocks.

03

Specification: βB−βW=−0.600, so the across-store relationship is twice as negative as the within relationship. CRE exposes this gap, but no Wald test is run here. βRE=−0.689372 is closer to βW while remaining a blend. Adding week effects is the minimum sensitivity analysis for common shocks; random slopes and residuals must be examined on real data.

13

Interpreting the results

  • βW=−0.600: in this synthetic dataset, +1 display point within the same store is associated with 0.6 fewer orders per 1,000 visits.
  • βB=−1.200 answers another question: comparing stores with persistently different average display shares.
  • βRE=−0.689 is neither βW nor βB; it combines both according to the estimated variance structure.

14

Supported and forbidden conclusions

Supported

  • Describe within and between associations separately.
  • Show which extra assumption standard RE adds.
  • Choose a specification aligned with the question and document uncertainty.

Forbidden

  • Turn Hausman into an automatic FE-otherwise-RE rule.
  • Present βW as causal without additional design and identification assumptions.
  • Believe fixed effects remove every source of endogeneity.

15

Possible marketing decision

For a within-store operating decision, prioritize βW and its uncertainty. If structural differences across stores also matter, report βB or a full CRE model. If causal interpretation is required, add an identification design; the panel model alone is insufficient.

16

When to use and when not to use

01

FE: within question, potentially correlated stable heterogeneity, time-invariant variables not central.

02

Standard RE: only when pooling within/between and orthogonality are defensible.

03

CRE/Mundlak: when both levels matter or their equality must be examined explicitly.

17

Reproducible implementations

Python 3.13 (standard library) and R 4.5 (base R) are the two reference implementations: same CSV, same declared formulas and same expected outputs. SPSS 31 and SAS 9.4 are secondary native snippets: they show FE, RE and CRE, but their RE estimators, variance components, covariances and degrees of freedom are not certified numerically equivalent. The βRE=−0.689372 value belongs to the reference moments method.

Download the Python reference script

Python 3.13 · reference

# Python 3.13, standard library only
python msc-p021-reference.py --csv msc-p021-panel.csv
# Expected: FE=-0.600000; RE=-0.689372; between=-1.200000

R 4.5 · reference

# R 4.5, base R only: exact companion to the Python reference
d <- read.csv("msc-p021-panel.csv")
G <- length(unique(d$store_id)); Tn <- 8
d$xbar <- ave(d$display_share_pct, d$store_id)
d$ybar <- ave(d$orders_per_1000, d$store_id)
d$xwithin <- d$display_share_pct - d$xbar
d$ywithin <- d$orders_per_1000 - d$ybar
pooled <- coef(lm(orders_per_1000 ~ display_share_pct, d))[2]
within <- sum(d$xwithin*d$ywithin)/sum(d$xwithin^2)
between_fit <- lm(ybar ~ xbar, unique(d[c("store_id","xbar","ybar")]))
between <- coef(between_fit)[2]
e <- d$ywithin - within*d$xwithin
sigma_e2 <- sum(e^2)/(G*(Tn-1)-1)
sigma_u2 <- max(0, sum(resid(between_fit)^2)/(G-2)-sigma_e2/Tn)
theta <- 1-sqrt(sigma_e2/(sigma_e2+Tn*sigma_u2))
re <- coef(lm(I(orders_per_1000-theta*ybar) ~ I(display_share_pct-theta*xbar), d))[2]
scores <- tapply(d$xwithin*e, d$store_id, sum)
cluster_se <- sqrt((G/(G-1))*sum(scores^2)/sum(d$xwithin^2)^2)
ci <- within + c(-1,1)*qt(.975, df=G-1)*cluster_se
c(pooled=pooled, within=within, between=between, re=re,
  contextual=between-within, cluster_se=cluster_se, ci_low=ci[1], ci_high=ci[2])

IBM SPSS Statistics 31 · secondary

* IBM SPSS Statistics 31 — secondary native syntax, not numerically certified.
GET DATA /TYPE=TXT /FILE='msc-p021-panel.csv'
 /DELCASE=LINE /DELIMITERS=',' /FIRSTCASE=2
 /VARIABLES=store_id A3 week F2 display_share_pct F8.1 orders_per_1000 F8.1.
AGGREGATE OUTFILE=* MODE=ADDVARIABLES /BREAK=store_id
 /xbar=MEAN(display_share_pct).
COMPUTE xwithin=display_share_pct-xbar.
UNIANOVA orders_per_1000 BY store_id WITH display_share_pct
 /METHOD=SSTYPE(3) /DESIGN=store_id display_share_pct.
MIXED orders_per_1000 WITH display_share_pct
 /FIXED=INTERCEPT display_share_pct | SSTYPE(3)
 /RANDOM=INTERCEPT | SUBJECT(store_id) COVTYPE(VC) /METHOD=REML.
MIXED orders_per_1000 WITH xwithin xbar
 /FIXED=INTERCEPT xwithin xbar /RANDOM=INTERCEPT | SUBJECT(store_id) COVTYPE(VC).

SAS 9.4 · secondary

/* SAS 9.4 — secondary native syntax, not numerically certified. */
proc import datafile="msc-p021-panel.csv" out=panel dbms=csv replace; guessingrows=max; run;
proc sql; create table cre as select *, mean(display_share_pct) as xbar
 from panel group by store_id; quit;
data cre; set cre; xwithin=display_share_pct-xbar; run;
proc panel data=panel; id store_id week;
 model orders_per_1000=display_share_pct / fixone;
 model orders_per_1000=display_share_pct / ranone; run;
proc mixed data=cre method=reml;
 class store_id; model orders_per_1000=xwithin xbar / solution;
 random intercept / subject=store_id; run;

All four blocks use the same CSV and estimands, not necessarily the same native algorithm. No software validates exogeneity, absence of time-varying confounding or causal scope.

18

Expected final deliverable

  1. 01

    Table of units, periods, missingness and within variation.

  2. 02

    FE, standard RE and CRE results with estimands, covariance, intervals and diagnostics.

  3. 03

    Written conclusion: what is within, between, testable, untestable, supported and forbidden.

19

Scientific sources and evidence level

  1. Hausman (1978)Foundational paper, specification test

    Formalizes comparison between an estimator efficient under the null and one consistent under a broader alternative. It does not provide an automatic business rule.

  2. Bell, Fairbrother & Jones (2019)Methodological review and simulations

    Separates within and between, explains standard RE blending, random slopes and why Hausman should not decide alone.

  3. Wooldridge (2019)CRE theory for balanced and unbalanced panels

    Establishes the within-coefficient equivalence obtained by adding time averages and states exogeneity conditions.

All three full texts were verified and tied to specific claims. The dataset, code and numerical results are synthetic MSC creations, without external empirical validation.

Method connections

Parent territoryMarketing response models: shape, delay and saturationRequiresPredictive or causal regression: what are you trying to estimate?Compare withHow do you address price endogeneity?

Read next

MSC-H-002Marketing response models: shape, delay and saturationMSC-P-018Predictive or causal regression: what are you trying to estimate?MSC-P-020How do you address price endogeneity?MSC-P-033How do you validate a marketing forecast?