Research & Evidence

Search for a method

Search titles, questions, territories and MSC identifiers.

40 results
  1. MSC-P-001How do you turn a marketing claim into a testable question?Decision Science
  2. MSC-P-002Correlation or causality: what can an analysis actually support?Marketing Measurement
  3. MSC-P-003How should uncertainty in a marketing result be expressed?Decision Science
  4. MSC-P-004Statistical significance or effect size: which result should be interpreted?Decision Science
  5. MSC-P-005How do you measure a marketing construct that is not directly observable?Market Research
  6. MSC-P-006How do you design and validate a measurement scale?Market Research
  7. MSC-P-007Alpha or omega: how should scale reliability be assessed?Market Research
  8. MSC-P-009PCA, EFA or CFA: which method should you choose?Market Research
  9. MSC-P-010When should you run a marketing experiment?Marketing Measurement
  10. MSC-P-011How do you design an A/B test that actually estimates an effect?Marketing Measurement
  11. MSC-P-012How many observations does an experiment need?Decision Science
  12. MSC-P-013How do you measure campaign incrementality with a control group?Marketing Measurement
  13. MSC-P-017How do you detect selection, contamination and attrition in an experiment?Marketing Measurement
  14. MSC-P-018Predictive or causal regression: what are you trying to estimate?Marketing Models
  15. MSC-P-019How do you diagnose a marketing regression before interpreting it?Marketing Models
  16. MSC-P-022How do you estimate price elasticity and its uncertainty?Pricing Science
  17. MSC-P-026Logit vs Probit: how do you choose for purchase probability?Customer Science
  18. MSC-P-029Which customers have the highest probability of churn?Customer Science
  19. MSC-P-027TAM, UTAUT or UTAUT2: which framework should be used to study technology acceptance?Market Research
  20. MSC-H-001Measurement and causality: how can a marketing effect be established?Marketing Measurement
  21. MSC-H-002Marketing response models: shape, delay and saturationMarketing Models
  22. MSC-H-003Pricing science: connecting price, demand and contributionPricing Science
  23. MSC-H-004Customer and choice science: behavior, value and heterogeneityCustomer Science
  24. MSC-H-005Measurement science: building valid indicatorsMarket Research
  25. MSC-H-006Statistical decision methods: choose, quantify, validateDecision Science
  26. MSC-P-008How do you validate a marketing measurement scale?Market Research
  27. MSC-P-014How do you design a marketing geo experiment?Marketing Measurement
  28. MSC-P-015How do you estimate an effect with difference-in-differences?Marketing Measurement
  29. MSC-P-020How do you address price endogeneity?Pricing Science
  30. MSC-P-021Fixed or random effects: which panel model should you choose?Marketing Models
  31. MSC-P-023How do you estimate a demand function?Pricing Science
  32. MSC-P-024How do you simulate a price-volume-margin scenario?Pricing Science
  33. MSC-P-028How do you estimate CLV with BG/NBD and Gamma-Gamma?Customer Science
  34. MSC-P-030How do you analyze retention with a survival model?Customer Science
  35. MSC-P-031How do you build a useful customer segmentation?Customer Science
  36. MSC-P-032How do you test segmentation stability?Customer Science
  37. MSC-P-033How do you validate a marketing forecast?Decision Science
  38. MSC-P-034How do you build a Monte Carlo simulation for a marketing decision?Decision Science
  39. MSC-P-035How do you model saturation and adstock?Marketing Models
  40. MSC-P-039Which statistical test should you choose?Decision Science
All methods
METHOD DOSSIERMSC-P-019Regression and econometricsVerified scientific dossier

How do you diagnose a marketing regression before interpreting it?

Diagnosis examines functional form, residuals, heteroskedasticity, collinearity, influence, temporal dependence and out-of-sample stability before reading coefficients.

Scientific editorial team : Marketing Science Center

Direct answer

Determine which interpretations remain compatible with diagnostics.

Diagnosis examines functional form, residuals, heteroskedasticity, collinearity, influence, temporal dependence and out-of-sample stability before reading coefficients.

Shmueli, 2010Arnold et al., 2020

01

Operational summary

Scientific question and scope

How do you diagnose a marketing regression before interpreting it?

Supported

Determine which interpretations remain compatible with diagnostics.

Forbidden

Validate a model from R-squared alone.

02

Three reading levels

  1. 01

    Decision-makerConnect the result to a declared decision, useful threshold and error cost.

  2. 02

    PractitionerFix population, unit, horizon, available variables and analysis rule before calculation.

  3. 03

    AnalystReproduce the calculation, quantify uncertainty and document diagnostics, failures and sensitivities.

03

Concrete marketing situation

Diagnosis examines functional form, residuals, heteroskedasticity, collinearity, influence, temporal dependence and out-of-sample stability before reading coefficients.

04

Scientific question and scope

How do you diagnose a marketing regression before interpreting it?

Diagnostic evidence

Required data

observation × residual

Population, unit of analysis, origin date and horizon must be declared in the deliverable. Without them, the estimand silently changes.

05

Why a simple analysis can fail

  • Validate a model from R-squared alone.
  • Residual plots
  • Robust variance
  • Influence and holdout

06

Method intuition

The method does not automatically turn an association into evidence. It links a declared question to an estimand, specification, compatible data and a bounded interpretation rule.

Diagnostic evidence

07

Required data

Scientific symbol dictionary

r_i
Standardized residual for observation or cluster i under the declared fitted model. Unit: standard deviation · Type: number · Role: diagnostic
h_i
Leverage of observation or cluster i under the declared design matrix. Unit: dimensionless · Type: number in [0, 1] · Role: diagnostic
D_i
Cook's distance for observation or cluster i under the declared model. Unit: dimensionless · Type: nonnegative number · Role: diagnostic
RMSE_holdout
Root mean squared prediction error on observations excluded from fitting. Unit: outcome · Type: nonnegative number · Role: diagnostic

Exact sealed engine inputs

residual_mean
Value: 0.02 · Unit: outcome · Type: number · Data status: synthetic
residual_sd
Value: 1.20 · Unit: outcome · Type: number · Data status: synthetic
maximum_absolute_standardized_residual
Value: 2.40 · Unit: dimensionless · Type: number · Data status: synthetic
maximum_leverage
Value: 0.08 · Unit: dimensionless · Type: number · Data status: synthetic
maximum_cooks_distance
Value: 0.06 · Unit: dimensionless · Type: number · Data status: synthetic
sample_size
Value: 100 · Unit: observation · Type: integer · Data status: synthetic
parameter_count
Value: 5 · Unit: parameter · Type: integer · Data status: parameter
cluster_count
Value: 30 · Unit: cluster · Type: integer · Data status: synthetic
holdout_rmse
Value: 1.35 · Unit: outcome · Type: number · Data status: synthetic

08

Formal model

Formal model

Report residual_mean, residual_sd, maximum_absolute_standardized_residual, maximum_leverage, maximum_cooks_distance, cluster_count and holdout_rmse; infer influence only after case-deletion or cluster-deletion refit sensitivity

Evidence claims: MSC-P019-C01 · MSC-P019-C02 · MSC-P019-C03 · MSC-P019-C04 · MSC-P019-C05

Scientific question and scope

Diagnostic evidence

09

Declared calculation

  1. 01

    Lock the model, residual definition, cluster structure and holdout split before inspecting extrema.

  2. 02

    Report descriptive residual mean/SD, maximum absolute standardized residual, leverage, Cook's distance and holdout RMSE without universal cutoffs.

  3. 03

    Require case-deletion or cluster-deletion refitting and sensitivity of the decision quantity before calling a point influential.

10

Numerical example or application case

Synthetic illustration — teaching values, not observedsynthetic data or declared parameters; no real observations

Sealed inputs

  • residual_mean=0.02 [outcome]
  • residual_sd=1.20 [outcome]
  • maximum_absolute_standardized_residual=2.40 [dimensionless]
  • maximum_leverage=0.08 [dimensionless]
  • maximum_cooks_distance=0.06 [dimensionless]
  • sample_size=100 [observation]
  • parameter_count=5 [parameter]
  • cluster_count=30 [cluster]
  • holdout_rmse=1.35 [outcome]

Reproducible results

  • descriptive_max_standardized_residual=2.4
  • descriptive_max_leverage=0.08
  • descriptive_max_Cook_D=0.06
  • holdout_RMSE=1.35

Verified dossier: claims are linked to passage-level sources, the Python/R calculation is reproduced, limits are explicit, and independent scientific review is complete.

11

Validity assumptions

  • Are residuals, leverage and Cook's distance defined for the exact fitted model and dependence structure?
  • Was the holdout excluded from all model and specification selection?
  • Does deletion/refit sensitivity materially change the decision-relevant quantity?

12

Diagnostics and uncertainty

Diagnostics and uncertainty

Residual plots · Robust variance · Influence and holdout

Bounded diagnostic

Descriptive inspection only: there is no universal cutoff here, and influence requires case-deletion/refit sensitivity before any conclusion.

Explicit uncertainty contract · not_estimable_from_sealed_inputs

Method : Explicit non-estimability assessment against the sealed input schema.

Target : Uncertainty and decision sensitivity of regression diagnostics.

Engine evidence : diagnostic=descriptive_inspection_only_no_universal_cutoff_influence_conclusion_requires_case_deletion_refit_sensitivity

Interpretation : Sealed extrema and one holdout RMSE do not provide refit distributions; case- or cluster-deletion refits are required.

Stop when

Validate a model from R-squared alone.

13

Result interpretation

  • Determine which interpretations remain compatible with diagnostics.
  • The result is conditional on the declared population, horizon, specification and diagnostics. It must not be extended to another decision without new justification.

14

Supported and forbidden conclusions

Supported

Determine which interpretations remain compatible with diagnostics.

Forbidden

Validate a model from R-squared alone.

15

Possible marketing decision

  1. 01

    Determine which interpretations remain compatible with diagnostics.

  2. 02

    The result is conditional on the declared population, horizon, specification and diagnostics. It must not be extended to another decision without new justification.

16

When to use — when to stop

Use when

Determine which interpretations remain compatible with diagnostics.

observation × residual

Stop when

Validate a model from R-squared alone.

Methodological alternatives

  • Use robust regression and compare decision quantities under alternative specifications.
  • Use cluster-deletion or leave-one-group-out refitting when dependence is grouped.
  • Use graphical residual checks and domain-specific falsification tests instead of universal numeric cutoffs.

17

Reproducibility contract

Python and R are the executable references. SPSS and SAS remain secondary syntaxes until checked on the same data, specification and diagnostics.

The engine selects only this page’s explicit slice and branch. Its hash, outputs, and diagnostics remain sealed in the atomic dossier; no generic fallback branch is allowed.

Slice fingerprint: fe38d763c26218df98d157e806962f0ec5f9b9a4baeca89b39123cc1db00b769

Formal model

Report residual_mean, residual_sd, maximum_absolute_standardized_residual, maximum_leverage, maximum_cooks_distance, cluster_count and holdout_rmse; infer influence only after case-deletion or cluster-deletion refit sensitivity

Verified dossier: claims are linked to passage-level sources, the Python/R calculation is reproduced, limits are explicit, and independent scientific review is complete.

18

Expected final deliverable

  1. 01

    How do you diagnose a marketing regression before interpreting it?observation × residual

  2. 02

    r_i · h_i · D_i · RMSE_holdout

  3. 03

    Report residual_mean, residual_sd, maximum_absolute_standardized_residual, maximum_leverage, maximum_cooks_distance, cluster_count and holdout_rmse; infer influence only after case-deletion or cluster-deletion refit sensitivity

  4. 04

    Residual plots · Robust variance · Influence and holdout

  5. 05

    Determine which interpretations remain compatible with diagnostics. / Validate a model from R-squared alone.

19

Scientific sources and evidence status

Verified dossier: claims are linked to passage-level sources, the Python/R calculation is reproduced, limits are explicit, and independent scientific review is complete.

  1. MSC-P019-C01Predictive accuracy should be evaluated on observations excluded from model fitting.
  2. MSC-P019-C02Fitting a regression does not itself establish a causal design.
  3. MSC-P019-C03Influence diagnostics target observations or clusters under the declared model.
  4. MSC-P019-C04Deletion diagnostics depend on the observational unit being removed.
  5. MSC-P019-C05Cook's distance is an influence diagnostic, not a model-validity verdict.
  1. Shmueli, 2010
  2. Arnold et al., 2020
  3. Zhu, Ibrahim & Cho, 2012

Method connections

Parent territoryMarketing response models: shape, delay and saturation

Read next

MSC-P-018Predictive or causal regression: what are you trying to estimate?MSC-P-022How do you estimate price elasticity and its uncertainty?MSC-P-026Logit vs Probit: how do you choose for purchase probability?