Which statistical test should you choose?
Choice starts from the question, estimand, sampling design, variable type and dependencies. Naming the test is only the final step.
Scientific editorial team: Marketing Science Center
Direct answer
Select a coherent procedure and report effect, interval and assumptions.
Choice starts from the question, estimand, sampling design, variable type and dependencies. Naming the test is only the final step.
01 · PROTOCOL
Operational summary
Three procedures are put to the test on 2000 simulated datasets: Welch's test, the rank test, and the popular rule that tests normality first and then chooses. Under the two declared null hypotheses, all three hold their announced level of 0.05 — the two-stage rule rejects 0.060000 and 0.056000 of the time. Verdict: PRETEST_RULE_HOLDS_ITS_LEVEL. But the real gap lies elsewhere: on skewed data, Welch detects only 0.319500 of effects against 0.422500 for the rank test. The choice is paid for in power, not in level.
02 · PROTOCOL
Concrete marketing situation
A team has measured the time spent before a purchase decision in two groups of thirty people. The spreadsheet offers a Student test, a colleague advises a nonparametric test because “the data are not normal”, another suggests testing normality first and letting the answer decide. Three pieces of advice, three procedures, and no obvious way to settle the matter from this one file.
03 · PROTOCOL
Scientific question
For these three declared procedures, this nominal level of 0.05, these two declared null hypotheses and these 2000 replications, does each procedure reject as often as it claims? And if so, which one best detects a real effect? The question is not about the team’s file, which cannot settle anything on its own, but about the behaviour of the procedures themselves when the truth is known.
04 · PROTOCOL
Why the simple approach can fail
Looking at the file and comparing p-values settles nothing: on this file the three procedures report 0.012060, 0.010993 and 0.010993, and no decision would change. That is the usual trap — one believes one is choosing a test by looking at the data, when a single dataset cannot say whether a procedure keeps its promises. And testing normality to decide amounts to letting the data choose their own judge, which is not the same thing as a test at 5%.
05 · PROTOCOL
Method intuition
The method consists in turning the question around. Instead of asking which test suits these data, one manufactures thousands of datasets whose truth is known, applies each procedure to all of them, and counts. Under a null hypothesis, a procedure honest at 5% rejects about five times in a hundred: if it rejects more often, it is not a test at 5%, whatever its name. Under an alternative hypothesis, one counts again, and this time the number is called power. Two counts, two distinct properties.
06 · PROTOCOL
Required data
Frozen before any reading: the gapless unit identifier, the group restricted to two values, the measured outcome, the three procedures, the nominal level of 0.05, the level of the normality test, the three scenarios, the 2000 replications, the 25 units per group, the recursion that produces the draws and the tolerance on the level. The sealed worked file holds 60 units. The case is synthetic and declared as such.
07 · PROTOCOL
Formal model and symbols
Welch's test compares means by dividing their gap by a standard error that allows unequal variances, with Satterthwaite degrees of freedom. The rank test replaces each value by its rank in the combined sample, sums the ranks of one group and compares that sum with what it would be with no difference, through a normal approximation with a tie correction and no continuity correction. The declared normality test is Jarque-Bera's: it combines skewness and kurtosis into a statistic compared with a chi-square law on two degrees of freedom.
08 · PROTOCOL
Declared calculation
The declared computation runs seven steps: validate the schema and refuse any non-conforming file; run the three procedures on the worked file and record what the two-stage rule chose; for each scenario, reset the recursion to its seed; draw 2000 replications of two groups of 25 by Box-Muller, exponentiating for the skewed scenarios and applying the declared shift to the second group; apply the three procedures to each replication and note the rejections and the test chosen; divide by the number of replications; apply the declared tolerance to the two null scenarios, then the verdict rule to the two-stage rule alone.
09 · PROTOCOL
End-to-end numeric example
Synthetic illustration — teaching values, not observed
On the sealed worked file: 30 units per group, Welch gives a statistic of 2.640102 on 37.036660 degrees of freedom and a tail of 0.012060; the rank test gives 278.000000, a standardized value of −2.542921 and a tail of 0.010993; Jarque-Bera gives 54.410246 and 37.041217, so the two-stage rule chooses the rank test and reports 0.010993. Over the 2000 replications: under the normal null, 0.058500 for Welch, 0.053500 for the rank test, 0.060000 for the rule, which chose Welch 0.953500 of the time; under the skewed null, 0.049000, 0.053500 and 0.056000, the rule having chosen Welch 0.058000 of the time. The six checks pass. Under the skewed alternative, 0.319500, 0.422500 and 0.424500. Verdict: PRETEST_RULE_HOLDS_ITS_LEVEL.
10 · PROTOCOL
Validity assumptions
The simulation assumes that units are independent, that the two groups are exchangeable under the null, and that the declared laws — normal and lognormal — represent situations of interest. It also assumes that the declared generator produces draws close enough to independence for the counts to be interpretable. Three things are not testable here: the behaviour of the procedures under unequal variances and unequal group sizes, that of other normality tests, and whether the mean or the rank answers the question the team is actually asking.
11 · PROTOCOL
Diagnostics and uncertainty
The diagnostic fits in two readings. The level: the six rates under the null run from 0.049000 to 0.060000, all inside the tolerance of 0.0125 around 0.05, whose standard error is about 0.0049 for 2000 replications — the tolerance is therefore about two and a half standard errors. The power: under the skewed alternative, the gap between Welch and the rank test, 0.319500 against 0.422500, is far wider than anything separating the levels. And the share of Welch choices by the two-stage rule, 0.953500 on normal data against 0.058000 on skewed data, shows that this rule is not a third procedure but a switch.
12 · PROTOCOL
Robustness and alternatives
Alternatives declared before results: rerun the simulation with unequal variances and unequal group sizes, the configuration where the literature reports the largest distortions; replace Jarque-Bera by Shapiro-Wilk, the most common normality test; raise the number of replications to tighten the rates; or replace the two tests by a comparison of quantiles, which answers a different question from the mean or the rank. Each variant must be announced before reading and reported even if it overturns the verdict.
13 · PROTOCOL
Result interpretation
The result is not the expected condemnation. In the configurations declared here, the two-stage rule holds its level, and the peer-reviewed study sealed beside this dossier found the same while recalling that the procedure remains formally incorrect. Both statements are true together: conditioning the choice of test on the data breaks the logic of the test, and in these particular configurations the damage does not show up in the overall level. What does show up is the power lost when a test of means is applied to skewed data.
14 · PROTOCOL
Allowed conclusions
Allowed: choosing the test from the question, the sampling design and the outcome scale, before seeing the data; publishing the effect size and its interval beside the p-value; saying that on this file the three procedures agree; using a rank test when the distribution is visibly skewed and the question concerns a general shift rather than a mean; and rerunning this kind of simulation before adopting an automatic rule.
15 · PROTOCOL
Forbidden conclusions
Forbidden: concluding from this dossier that the two-stage rule is correct, when it remains formally faulty and was tested only on the declared configurations; generalizing these rates to unequal variances or unequal group sizes, which were not examined; choosing the test after seeing which procedure gives the lowest p-value; presenting a p-value alone as a result; or reading a failure to reject as evidence of no effect. Also forbidden: presenting this synthetic case as an observed measurement.
16 · PROTOCOL
Possible marketing decision
The reasonable decision is to write the choice of test into the protocol, before collection, from the question asked and the sampling design — and to stick to it. If the expected distribution is skewed, that leads here to the rank test, which detects a third more effects. If the team insists on a difference of means because it is the mean that carries economic value, it keeps Welch and accepts the lower power knowingly. What no automatic rule will do in its place is decide which of the two questions it is asking.
17 · PROTOCOL
When to use or avoid the method
Use this approach before fixing an analysis procedure in a protocol or in a tool. Avoid it as a way of choosing a test after the fact, which is precisely what it forbids. Avoid it too when units are not independent — repeated measures, clusters, panels — since none of the three procedures examined here accounts for that. And avoid it when the question concerns a causal effect: the choice of test does not replace a design that identifies that effect.
18 · PROTOCOL
Implementations and final deliverable
The CC0 CSV holds the 60 synthetic units of the worked file. Python and R are the reference implementations; the SPSS block is not a rewrite but is generated from the reference's own computation, so the two cannot drift; SAS reproduces the declared recursion and implements the rank test in a data step, but takes its distribution tails from its own functions. The final deliverable gathers the data, the variable dictionary, the three declared procedures, the nominal level, the three scenarios, the recursion and its exact form in double arithmetic, the tolerance, the seven numbered steps, the results on the worked file, the rejection rates by scenario, the share of each test chosen, the six flags, the verdict and the software versions.
Synthetic data · CC0
msc-p039-two-group-outcome.csv ↓Reproducibility protocol
msc-p039-reproducibility-readme.md ↓Python reference · MIT
msc-p039-reference.py ↓R reference · MIT
msc-p039-reference.R ↓SPSS implementation · MIT
msc-p039-secondary.sps ↓SAS implementation · MIT
msc-p039-secondary.sas ↓19 · PROTOCOL
Sources and evidence level
Rochon, Gondan and Kieser describe exactly the rule examined here — the outcome of the preliminary test determines the method used next — find that this preliminary test seriously alters the conditional error, conclude that the two-stage procedure may be judged incorrect from a formal perspective, and nevertheless observe that it seemed to maintain the nominal level satisfactorily. Greenland and co-authors recall that every method of inference rests on a complex web of assumptions, and that a very small p-value does not tell which one is wrong. Lakens recalls that the effect size is the most important outcome of an empirical study. Rousselet, Pernet and Wilcox recall that summarising a distribution by its mean is a choice, not an obvious step. Founding references: Rochon, Gondan and Kieser (2012, BMC Medical Research Methodology, CC BY open access, peer-reviewed) and Greenland et al. (2016, European Journal of Epidemiology, open access, peer-reviewed). Recent developments: Lakens (2013, Frontiers in Psychology, CC BY open access, peer-reviewed) and Rousselet, Pernet and Wilcox (2017, accepted manuscript deposited at the University of Edinburgh; published in the European Journal of Neuroscience, peer-reviewed).
- Rochon, Gondan & Kieser (2012) Full text verified, with a short excerpt located in the source.
- Greenland et al. (2016) Full text verified, with a short excerpt located in the source.
- Lakens (2013) Full text verified, with a short excerpt located in the source.
- Rousselet, Pernet & Wilcox (2017) Full text verified, with a short excerpt located in the source.
Method connections

