Research & Evidence

Search for a method

Search titles, questions, territories and MSC identifiers.

41 results
  1. MSC-P-001How do you turn a marketing claim into a testable question?Decision Science↗
  2. MSC-P-002Correlation or causality: what can an analysis actually support?Marketing Measurement↗
  3. MSC-P-003How should uncertainty in a marketing result be expressed?Decision Science↗
  4. MSC-P-004Statistical significance or effect size: which result should be interpreted?Decision Science↗
  5. MSC-P-005How do you measure a marketing construct that is not directly observable?Market Research↗
  6. MSC-P-006How do you design and validate a measurement scale?Market Research↗
  7. MSC-P-007Alpha or omega: how should scale reliability be assessed?Market Research↗
  8. MSC-P-009PCA, EFA or CFA: which method should you choose?Market Research↗
  9. MSC-P-010When should you run a marketing experiment?Marketing Measurement↗
  10. MSC-P-011How do you design an A/B test that actually estimates an effect?Marketing Measurement↗
  11. MSC-P-012How many observations does an experiment need?Decision Science↗
  12. MSC-P-013How do you measure campaign incrementality with a control group?Marketing Measurement↗
  13. MSC-P-017How do you detect selection, contamination and attrition in an experiment?Marketing Measurement↗
  14. MSC-P-018Predictive or causal regression: what are you trying to estimate?Marketing Models↗
  15. MSC-P-019How do you diagnose a marketing regression before interpreting it?Marketing Models↗
  16. MSC-P-022How do you estimate price elasticity and its uncertainty?Pricing Science↗
  17. MSC-P-026Logit vs Probit: how do you choose for purchase probability?Customer Science↗
  18. MSC-P-029Which customers have the highest probability of churn?Customer Science↗
  19. MSC-P-027TAM, UTAUT or UTAUT2: which framework should be used to study technology acceptance?Market Research↗
  20. MSC-H-001Measurement and causality: how can a marketing effect be established?Marketing Measurement↗
  21. MSC-H-002Marketing response models: shape, delay and saturationMarketing Models↗
  22. MSC-H-003Pricing science: connecting price, demand and contributionPricing Science↗
  23. MSC-H-004Customer and choice science: behavior, value and heterogeneityCustomer Science↗
  24. MSC-H-005Measurement science: building valid indicatorsMarket Research↗
  25. MSC-H-006Statistical decision methods: choose, quantify, validateDecision Science↗
  26. MSC-P-008How do you validate a marketing measurement scale?Market Research↗
  27. MSC-P-014How do you design a marketing geo experiment?Marketing Measurement↗
  28. MSC-P-015How do you estimate an effect with difference-in-differences?Marketing Measurement↗
  29. MSC-P-020How do you address price endogeneity?Pricing Science↗
  30. MSC-P-021Fixed or random effects: which panel model should you choose?Marketing Models↗
  31. MSC-P-023How do you estimate a demand function?Pricing Science↗
  32. MSC-P-024How do you simulate a price-volume-margin scenario?Pricing Science↗
  33. MSC-P-028How do you estimate CLV with BG/NBD and Gamma-Gamma?Customer Science↗
  34. MSC-P-030How do you analyze retention with a survival model?Customer Science↗
  35. MSC-P-043What is a subscriber worth with only six months of retention data?Customer Science↗
  36. MSC-P-031How do you build a useful customer segmentation?Customer Science↗
  37. MSC-P-032How do you test segmentation stability?Customer Science↗
  38. MSC-P-033How do you validate a marketing forecast?Decision Science↗
  39. MSC-P-034How do you build a Monte Carlo simulation for a marketing decision?Decision Science↗
  40. MSC-P-035How do you model saturation and adstock?Marketing Models↗
  41. MSC-P-039Which statistical test should you choose?Decision Science↗
← All methods
METHOD DOSSIERMSC-P-039Evidence foundationsVerified scientific dossier

Which statistical test should you choose?

Choice starts from the question, estimand, sampling design, variable type and dependencies. Naming the test is only the final step.

Scientific editorial team: Marketing Science Center

Direct answer

Select a coherent procedure and report effect, interval and assumptions.

Choice starts from the question, estimand, sampling design, variable type and dependencies. Naming the test is only the final step.

Welch, 1947ASA, 2016

01 · PROTOCOL

Operational summary

Three procedures are put to the test on 2000 simulated datasets: Welch's test, the rank test, and the popular rule that tests normality first and then chooses. Under the two declared null hypotheses, all three hold their announced level of 0.05 — the two-stage rule rejects 0.060000 and 0.056000 of the time. Verdict: PRETEST_RULE_HOLDS_ITS_LEVEL. But the real gap lies elsewhere: on skewed data, Welch detects only 0.319500 of effects against 0.422500 for the rank test. The choice is paid for in power, not in level.

02 · PROTOCOL

Concrete marketing situation

A team has measured the time spent before a purchase decision in two groups of thirty people. The spreadsheet offers a Student test, a colleague advises a nonparametric test because “the data are not normal”, another suggests testing normality first and letting the answer decide. Three pieces of advice, three procedures, and no obvious way to settle the matter from this one file.

03 · PROTOCOL

Scientific question

For these three declared procedures, this nominal level of 0.05, these two declared null hypotheses and these 2000 replications, does each procedure reject as often as it claims? And if so, which one best detects a real effect? The question is not about the team’s file, which cannot settle anything on its own, but about the behaviour of the procedures themselves when the truth is known.

04 · PROTOCOL

Why the simple approach can fail

Looking at the file and comparing p-values settles nothing: on this file the three procedures report 0.012060, 0.010993 and 0.010993, and no decision would change. That is the usual trap — one believes one is choosing a test by looking at the data, when a single dataset cannot say whether a procedure keeps its promises. And testing normality to decide amounts to letting the data choose their own judge, which is not the same thing as a test at 5%.

05 · PROTOCOL

Method intuition

The method consists in turning the question around. Instead of asking which test suits these data, one manufactures thousands of datasets whose truth is known, applies each procedure to all of them, and counts. Under a null hypothesis, a procedure honest at 5% rejects about five times in a hundred: if it rejects more often, it is not a test at 5%, whatever its name. Under an alternative hypothesis, one counts again, and this time the number is called power. Two counts, two distinct properties.

06 · PROTOCOL

Required data

Frozen before any reading: the gapless unit identifier, the group restricted to two values, the measured outcome, the three procedures, the nominal level of 0.05, the level of the normality test, the three scenarios, the 2000 replications, the 25 units per group, the recursion that produces the draws and the tolerance on the level. The sealed worked file holds 60 units. The case is synthetic and declared as such.

07 · PROTOCOL

Formal model and symbols

Welch's test compares means by dividing their gap by a standard error that allows unequal variances, with Satterthwaite degrees of freedom. The rank test replaces each value by its rank in the combined sample, sums the ranks of one group and compares that sum with what it would be with no difference, through a normal approximation with a tie correction and no continuity correction. The declared normality test is Jarque-Bera's: it combines skewness and kurtosis into a statistic compared with a chi-square law on two degrees of freedom.

08 · PROTOCOL

Declared calculation

The declared computation runs seven steps: validate the schema and refuse any non-conforming file; run the three procedures on the worked file and record what the two-stage rule chose; for each scenario, reset the recursion to its seed; draw 2000 replications of two groups of 25 by Box-Muller, exponentiating for the skewed scenarios and applying the declared shift to the second group; apply the three procedures to each replication and note the rejections and the test chosen; divide by the number of replications; apply the declared tolerance to the two null scenarios, then the verdict rule to the two-stage rule alone.

09 · PROTOCOL

End-to-end numeric example

Synthetic illustration — teaching values, not observed

On the sealed worked file: 30 units per group, Welch gives a statistic of 2.640102 on 37.036660 degrees of freedom and a tail of 0.012060; the rank test gives 278.000000, a standardized value of −2.542921 and a tail of 0.010993; Jarque-Bera gives 54.410246 and 37.041217, so the two-stage rule chooses the rank test and reports 0.010993. Over the 2000 replications: under the normal null, 0.058500 for Welch, 0.053500 for the rank test, 0.060000 for the rule, which chose Welch 0.953500 of the time; under the skewed null, 0.049000, 0.053500 and 0.056000, the rule having chosen Welch 0.058000 of the time. The six checks pass. Under the skewed alternative, 0.319500, 0.422500 and 0.424500. Verdict: PRETEST_RULE_HOLDS_ITS_LEVEL.

10 · PROTOCOL

Validity assumptions

The simulation assumes that units are independent, that the two groups are exchangeable under the null, and that the declared laws — normal and lognormal — represent situations of interest. It also assumes that the declared generator produces draws close enough to independence for the counts to be interpretable. Three things are not testable here: the behaviour of the procedures under unequal variances and unequal group sizes, that of other normality tests, and whether the mean or the rank answers the question the team is actually asking.

11 · PROTOCOL

Diagnostics and uncertainty

The diagnostic fits in two readings. The level: the six rates under the null run from 0.049000 to 0.060000, all inside the tolerance of 0.0125 around 0.05, whose standard error is about 0.0049 for 2000 replications — the tolerance is therefore about two and a half standard errors. The power: under the skewed alternative, the gap between Welch and the rank test, 0.319500 against 0.422500, is far wider than anything separating the levels. And the share of Welch choices by the two-stage rule, 0.953500 on normal data against 0.058000 on skewed data, shows that this rule is not a third procedure but a switch.

12 · PROTOCOL

Robustness and alternatives

Alternatives declared before results: rerun the simulation with unequal variances and unequal group sizes, the configuration where the literature reports the largest distortions; replace Jarque-Bera by Shapiro-Wilk, the most common normality test; raise the number of replications to tighten the rates; or replace the two tests by a comparison of quantiles, which answers a different question from the mean or the rank. Each variant must be announced before reading and reported even if it overturns the verdict.

13 · PROTOCOL

Result interpretation

The result is not the expected condemnation. In the configurations declared here, the two-stage rule holds its level, and the peer-reviewed study sealed beside this dossier found the same while recalling that the procedure remains formally incorrect. Both statements are true together: conditioning the choice of test on the data breaks the logic of the test, and in these particular configurations the damage does not show up in the overall level. What does show up is the power lost when a test of means is applied to skewed data.

14 · PROTOCOL

Allowed conclusions

Allowed: choosing the test from the question, the sampling design and the outcome scale, before seeing the data; publishing the effect size and its interval beside the p-value; saying that on this file the three procedures agree; using a rank test when the distribution is visibly skewed and the question concerns a general shift rather than a mean; and rerunning this kind of simulation before adopting an automatic rule.

15 · PROTOCOL

Forbidden conclusions

Forbidden: concluding from this dossier that the two-stage rule is correct, when it remains formally faulty and was tested only on the declared configurations; generalizing these rates to unequal variances or unequal group sizes, which were not examined; choosing the test after seeing which procedure gives the lowest p-value; presenting a p-value alone as a result; or reading a failure to reject as evidence of no effect. Also forbidden: presenting this synthetic case as an observed measurement.

16 · PROTOCOL

Possible marketing decision

The reasonable decision is to write the choice of test into the protocol, before collection, from the question asked and the sampling design — and to stick to it. If the expected distribution is skewed, that leads here to the rank test, which detects a third more effects. If the team insists on a difference of means because it is the mean that carries economic value, it keeps Welch and accepts the lower power knowingly. What no automatic rule will do in its place is decide which of the two questions it is asking.

17 · PROTOCOL

When to use or avoid the method

Use this approach before fixing an analysis procedure in a protocol or in a tool. Avoid it as a way of choosing a test after the fact, which is precisely what it forbids. Avoid it too when units are not independent — repeated measures, clusters, panels — since none of the three procedures examined here accounts for that. And avoid it when the question concerns a causal effect: the choice of test does not replace a design that identifies that effect.

18 · PROTOCOL

Implementations and final deliverable

The CC0 CSV holds the 60 synthetic units of the worked file. Python and R are the reference implementations; the SPSS block is not a rewrite but is generated from the reference's own computation, so the two cannot drift; SAS reproduces the declared recursion and implements the rank test in a data step, but takes its distribution tails from its own functions. The final deliverable gathers the data, the variable dictionary, the three declared procedures, the nominal level, the three scenarios, the recursion and its exact form in double arithmetic, the tolerance, the seven numbered steps, the results on the worked file, the rejection rates by scenario, the share of each test chosen, the six flags, the verdict and the software versions.

19 · PROTOCOL

Sources and evidence level

Rochon, Gondan and Kieser describe exactly the rule examined here — the outcome of the preliminary test determines the method used next — find that this preliminary test seriously alters the conditional error, conclude that the two-stage procedure may be judged incorrect from a formal perspective, and nevertheless observe that it seemed to maintain the nominal level satisfactorily. Greenland and co-authors recall that every method of inference rests on a complex web of assumptions, and that a very small p-value does not tell which one is wrong. Lakens recalls that the effect size is the most important outcome of an empirical study. Rousselet, Pernet and Wilcox recall that summarising a distribution by its mean is a choice, not an obvious step. Founding references: Rochon, Gondan and Kieser (2012, BMC Medical Research Methodology, CC BY open access, peer-reviewed) and Greenland et al. (2016, European Journal of Epidemiology, open access, peer-reviewed). Recent developments: Lakens (2013, Frontiers in Psychology, CC BY open access, peer-reviewed) and Rousselet, Pernet and Wilcox (2017, accepted manuscript deposited at the University of Edinburgh; published in the European Journal of Neuroscience, peer-reviewed).

  1. Rochon, Gondan & Kieser (2012) Full text verified, with a short excerpt located in the source.
  2. Greenland et al. (2016) Full text verified, with a short excerpt located in the source.
  3. Lakens (2013) Full text verified, with a short excerpt located in the source.
  4. Rousselet, Pernet & Wilcox (2017) Full text verified, with a short excerpt located in the source.

Method connections

Parent territoryStatistical decision methods: choose, quantify, validateRequiresStatistical significance or effect size: which result should be interpreted?Compare withHow many observations does an experiment need?

Read next

MSC-H-006Statistical decision methods: choose, quantify, validate→MSC-P-003How should uncertainty in a marketing result be expressed?→MSC-P-004Statistical significance or effect size: which result should be interpreted?→MSC-P-012How many observations does an experiment need?→