# MSC-P-039 — choosing a statistical test reproducibility protocol

## Frozen design

- Synthetic data only; no real user or session data.
- A worked file of 60 sampling units generated with seed `20260915`: two groups of thirty, a right-skewed outcome, the treated group shifted multiplicatively. It exists so that the three candidate procedures can be run once on a concrete case and be seen to disagree. **One dataset settles nothing**, which is the reason the simulation follows.
- Three candidate procedures, declared before any result is read:
  - **Welch**, the two-sample test with unequal variances and Satterthwaite degrees of freedom;
  - **the rank test**, Mann-Whitney's U through its normal approximation with the declared tie correction and **no continuity correction**;
  - **the two-stage rule**, the popular one: a Jarque-Bera normality test at level 0.05 in each group, then Welch if both pass and the rank test otherwise.
- Declared simulation: 2000 replications, 25 units per group, nominal level 0.05, and three scenarios — a normal null, a skewed null (lognormal with log standard deviation 0.85), and a skewed alternative with a log shift of 0.45. Under the two nulls the two groups are drawn from the same distribution.
- Declared randomness: the linear congruential recursion `state ← (1103515245 · state + 12345) mod 2^31` seeded at `20260915`, transformed to normal draws by Box-Muller. Written directly the product reaches 2.4e18, above 2^53, so every implementation splits the multiplier as `16838 · 65536 + 20077`, which is exact arithmetic and not an approximation.
- Declared threshold: under each null, a procedure's rejection rate must lie within **0.0125** of the nominal 0.05. With 2000 replications the standard error of a rate near 0.05 is about 0.0049, so the tolerance is about two and a half standard errors.
- The verdict concerns only the two-stage rule and is fail-closed: it declares the rule to hold its level only when it does so under **both** nulls.
- Power under the skewed alternative is reported without a threshold. It is not a level property and cannot be judged against 0.05.

## Expected output

`control_units=30`, `treated_units=30`,
`welch_statistic=2.640102`, `welch_degrees_of_freedom=37.036660`, `welch_probability=0.012060`,
`rank_statistic=278.000000`, `rank_standardized=-2.542921`, `rank_probability=0.010993`,
`control_normality_statistic=54.410246`, `control_normality_probability=0.000000`,
`treated_normality_statistic=37.041217`, `treated_normality_probability=0.000000`,
`pretest_choice=mann-whitney`, `pretest_probability=0.010993`,
`replications=2000`, `units_per_group=25`, `nominal_level=0.05`,
`normal_null_welch=0.058500`, `normal_null_mann_whitney=0.053500`, `normal_null_pretest=0.060000`, `normal_null_pretest_chose_welch=0.953500`,
`skewed_null_welch=0.049000`, `skewed_null_mann_whitney=0.053500`, `skewed_null_pretest=0.056000`, `skewed_null_pretest_chose_welch=0.058000`,
`skewed_alternative_welch=0.319500`, `skewed_alternative_mann_whitney=0.422500`, `skewed_alternative_pretest=0.424500`, `skewed_alternative_pretest_chose_welch=0.058000`,
all six `_flag` lines `PASS`, `level_tolerance=0.0125`,
`verdict=PRETEST_RULE_HOLDS_ITS_LEVEL`.

## Run

```text
python msc-p039-reference.py msc-p039-two-group-outcome.csv
Rscript msc-p039-reference.R msc-p039-two-group-outcome.csv
```

The SPSS syntax carries its computation in a `BEGIN PROGRAM Python3` block. That block is not a re-implementation: it is generated from the reference's own computation, so the two cannot drift, and it was executed outside SPSS and printed output identical to the Python reference, line for line. The R file uses base R's `pt()` and `pnorm()` where the Python reference computes the same tails from the incomplete beta and the error function; the two agree far beyond six decimals. The SAS program reproduces the declared recursion and implements the rank test in the data step, but takes its tail probabilities from PROBT and PROBNORM; it checks the flags and the verdict rather than every digit. The SPSS, SAS and R runtimes are not available in the local verification environment, so those three files are reviewed statically and must be run independently before relying on their printed output.

## Variable dictionary

| Column | Type | Unit | Definition |
|---|---|---|---|
| `unit_id` | text, `U001`…`U060` | — | Sampling unit identifier, ordered without gaps. |
| `group` | text, `control` or `treated` | — | Declared group of the sampling unit. |
| `minutes` | number > 0 | minutes | Time observed before the purchase decision. |

Derived quantities printed by the reference scripts:

| Output | Definition |
|---|---|
| `welch_statistic`, `welch_degrees_of_freedom`, `welch_probability` | Welch's test on the worked file: the standardized difference of means, its Satterthwaite degrees of freedom, and the two-sided tail. |
| `rank_statistic`, `rank_standardized`, `rank_probability` | Mann-Whitney's U computed on the control group, its standardized value, and the two-sided normal tail. The standardized value is negative because the control group holds the lower ranks. |
| `control_normality_statistic`, `treated_normality_statistic` and their probabilities | Jarque-Bera in each group. On this file both statistics are large and both tails print as `0.000000`. |
| `pretest_choice`, `pretest_probability` | Which test the two-stage rule selects on this file, and the tail it therefore reports. |
| `<scenario>_welch`, `<scenario>_mann_whitney`, `<scenario>_pretest` | Share of the 2000 replications in which each procedure rejects at 0.05. Under a null this is the procedure's actual level; under the alternative it is its power. |
| `<scenario>_pretest_chose_welch` | Share of replications in which the two-stage rule selected Welch. It shows that the rule is a switch, not a third procedure. |
| `<null>_<procedure>_flag` | `PASS` when that procedure's rejection rate under that null lies within 0.0125 of 0.05. |
| `verdict` | `PRETEST_RULE_HOLDS_ITS_LEVEL` only when the two-stage rule passes under both nulls; otherwise `PRETEST_RULE_DOES_NOT_HOLD_ITS_LEVEL`. |

## Numbered analysis steps

1. Load the CSV; refuse any file whose header, cell count, unit order, group labels or positivity differ from the frozen schema, or that leaves either group below ten units.
2. Run Welch, the rank test and Jarque-Bera on the worked file, and record which test the two-stage rule selects and what it reports.
3. For each of the three declared scenarios, reset the declared recursion to its seed.
4. Draw 2000 replications of two groups of 25 by Box-Muller over the declared recursion, exponentiating for the skewed scenarios and applying the declared shift to the second group.
5. In each replication, apply all three procedures and record whether each rejects at 0.05, and which test the two-stage rule chose.
6. Divide by the number of replications to obtain each procedure's rejection share in each scenario.
7. Apply the declared tolerance to the two null scenarios to obtain the six flags, apply the fail-closed verdict rule to the two-stage rule, and print every quantity of the expected output; a third party compares the printed lines with the expected output above, digit for digit.

## Scientific boundary

The worked file shows the three procedures reporting 0.012060, 0.010993 and 0.010993 — close enough that no decision would change, which is precisely why a single dataset cannot arbitrate between them. Only the simulation can, and only for the configurations it declares.

On those configurations the answer is not the cautionary tale one might expect. All three procedures hold their nominal level: the two-stage rule rejects 6.00 percent of the time under the normal null and 5.60 percent under the skewed null, both inside the declared tolerance. That agrees with the published study sealed alongside this dossier, which found the same and stated plainly that the procedure is nonetheless formally incorrect. Both statements are true at once: conditioning the choice of test on the data breaks the logic of the test, and in these particular configurations the damage does not show up in the overall level.

What the simulation does show sharply is that the choice matters for **power**, not for level. Under the skewed alternative, Welch rejects 31.95 percent of the time and the rank test 42.25 percent — a third more. The two-stage rule reaches 42.45 percent, but only because it selected Welch in 5.8 percent of skewed replications and 95.35 percent of normal ones: it is not a third procedure, it is a switch that lands on whichever test the data resemble. The honest reason to think about which test to use is that the wrong one wastes real evidence.

Four limits follow. The configurations most often reported as problematic — unequal variances, unequal group sizes, small samples — are outside this declared design and were not examined; adding them after seeing these results would be the scenario-hunting this dossier argues against, so they are named here as the next step instead. The normality pretest used is Jarque-Bera, chosen because it is computable exactly in every target language; Shapiro-Wilk, the more common choice, may behave differently. The declared recursion is a simple congruential generator chosen for cross-language reproducibility rather than for generator quality, and the whole simulation was rerun independently with a Mersenne Twister to check that the conclusions do not depend on it. And nothing here concerns whether the estimand is the right one: a test that holds its level can still answer the wrong question.

Dataset license: CC0-1.0. Code license: MIT.
