StatisticsSciPy 1.18 · pandas 3.0 · statsmodels 0.15 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
57 small wins to finish your pathNext lesson →

Choosing a statistical test

Choosing a statistical test is the step of matching a question to the test whose assumptions fit it, decided by the type of data, the number of groups, whether the groups are paired and what is known about the population.

Last updated: 07 Oct, 2026 · SciPy 1.18

The video's interview advice is to know why a test is used. This lesson puts the tests from the previous lessons side by side, adds the rank-based alternatives, and covers two things a p-value does not tell you: how big the effect is, and what happens when you run many tests.

Asking four questions about the data

  • What type is the variable? Numeric values call for tests on means; categories call for tests on counts and proportions (see Types of variables).
  • How many groups? One group against a stated value, two groups, or three and more.
  • Are the groups independent or paired? Different subjects, or the same subjects measured twice.
  • What is known or assumed? A known σ, roughly normal data, a large enough sample, expected counts of at least 5.
A decision chart. Means of a numeric variable: one mean with sigma known is a z-test, sigma unknown a one-sample t-test, two independent groups a Welch t-test or Mann-Whitney U, two paired measurements a paired t-test or Wilcoxon, three or more groups one-way ANOVA or Kruskal-Wallis. Counts of a categorical variable: one proportion is a z-test for a proportion, one variable against expected shares a chi-square goodness-of-fit test, two categorical variables a chi-square test of independence, a small 2 by 2 table Fisher's exact test. Two numeric variables: Pearson r, Spearman, or covariance.

Matching each test to its example

QuestionDataTestSciPy / statsmodels
Did the medication change mean IQ? (σ = 15 known)numeric, 1 groupOne-sample z-testby hand with norm
Same, with only the sample s = 20numeric, 1 groupOne-sample t-test and the t distributionttest_1samp
Do class A and class B have the same mean age?numeric, 2 independent groupsTwo-sample and paired t-teststtest_ind(equal_var=False)
Did weights change between week 10 and week 20?numeric, pairedPaired t-testttest_rel
Do 70% of residents own a phone?yes/no, 1 groupZ-test for a proportionproportions_ztest, binomtest
Has the age mix changed since 2000?1 categorical variableChi-square goodness-of-fit testchisquare
Is smoking related to sex?2 categorical variablesChi-square test of independencechi2_contingency
Do 3 iris species share a mean petal width?numeric, 3 groupsOne-way ANOVA (F-test)f_oneway

Choosing parametric or non-parametric tests

A parametric test assumes a form for the population, usually normal, and tests a parameter such as the mean: z, t and ANOVA. A non-parametric test makes no such assumption. The rank-based ones replace the values by their ranks, so outliers and skew matter less; they test whether one group tends to have larger values than another. Use them for small, skewed or ordinal data, and for data with outliers.

Parametric testRank-based alternativeSciPy
One-sample or paired t-testWilcoxon signed-rank testwilcoxon
Two-sample t-testMann-Whitney U testmannwhitneyu
One-way ANOVAKruskal-Wallis H testkruskal
Pearson correlationSpearman rank correlationspearmanr

On the notebook's class ages both kinds of test agree:

ExampleFrom the video's notebook, run on SciPy 1.18.1
import numpy as np
from scipy import stats

np.random.seed(6)
school_ages = stats.poisson.rvs(loc=18, mu=35, size=1500)
classA_ages = stats.poisson.rvs(loc=18, mu=30, size=60)
np.random.seed(12)
ClassB_ages = stats.poisson.rvs(loc=18, mu=33, size=60)
print("Welch t-test p:  ", round(stats.ttest_ind(classA_ages, ClassB_ages, equal_var=False).pvalue, 6))
print("Mann-Whitney U p:", round(stats.mannwhitneyu(classA_ages, ClassB_ages).pvalue, 6))

The p-values differ in size but both are below 0.05, so both reject H₀. When the data are close to normal, the two kinds of test usually agree; they part ways when outliers or heavy skew are present.

Measuring the effect size

A p-value says whether a difference is likely to be more than sampling noise. It does not say how big the difference is. An effect size does: Cohen's d is the difference in means in units of the standard deviation, and Cramér's V runs from 0 (no association) to 1 for a contingency table.

Cohen's d for one sample and Cramér's V for an r × c table
ExampleRun on SciPy 1.18.1
import numpy as np
import pandas as pd
import seaborn as sns
from scipy import stats

d_iq = (140 - 100) / 15                                   # Cohen's d, the z-test example
x = np.array([88, 92, 94, 94, 96, 97, 97, 97, 99, 99,
              105, 109, 109, 109, 110, 112, 112, 113, 114, 115])
d_20 = (x.mean() - 100) / x.std(ddof=1)                   # Cohen's d, the 20 patients
tips = sns.load_dataset("tips")
table = pd.crosstab(tips["sex"], tips["smoker"])
chi2 = stats.chi2_contingency(table, correction=False).statistic
cramers_v = np.sqrt(chi2 / (table.values.sum() * (min(table.shape) - 1)))
print("Cohen's d, IQ 140 vs 100:", round(d_iq, 2))
print("Cohen's d, 20 patients:  ", round(d_20, 2))
print("Cramer's V, sex by smoker:", round(cramers_v, 4))

A d of 2.67 is a huge effect: the medication group sits 2.67 population standard deviations above the norm. The 20 patients' d of 0.36 is small to moderate, and Cramér's V of 0.0028 for sex by smoking is essentially no association. The opposite also happens: with a very large sample a tiny difference becomes "significant".

ExampleRun on SciPy 1.18.1
import numpy as np
from scipy import stats

rng = np.random.default_rng(42)
a = rng.normal(loc=100.0, scale=15, size=200_000)
b = rng.normal(loc=100.3, scale=15, size=200_000)      # a 0.3-point true difference
r = stats.ttest_ind(a, b, equal_var=False)
d = (b.mean() - a.mean()) / np.sqrt((a.var(ddof=1) + b.var(ddof=1)) / 2)
print(f"p = {r.pvalue:.2g}, Cohen's d = {d:.3f}")

With 200,000 people per group, a 0.3-point IQ difference gives a p-value far below 0.05, but d is about 0.02, too small to matter for any decision. Report the effect size next to the p-value.

Correcting for multiple tests

Every test at α = 0.05 has a 5% chance of a Type I error when H₀ is true (see Type I and Type II errors). Run 20 independent tests on data with no real effects and the chance of at least one false alarm is 1 − 0.95²⁰ = 0.642. The Bonferroni correction divides α by the number of tests m, so each test uses α/m and the chance of any false alarm stays at most α.

ExampleRun on SciPy 1.18.1
import numpy as np
from scipy import stats

alpha, m = 0.05, 20
print("P(at least one false alarm in 20 tests):", round(1 - (1 - alpha) ** m, 3))
rng = np.random.default_rng(42)
p_values = [stats.ttest_ind(rng.normal(size=30), rng.normal(size=30)).pvalue for _ in range(m)]
p_values = np.array(p_values)                            # H0 is true in every test
print("tests with p <= 0.05:         ", int((p_values <= alpha).sum()))
print("tests with p <= 0.05/20:      ", int((p_values <= alpha / m).sum()))
print("smallest p:", round(p_values.min(), 4))

In this run of 20 tests where H₀ is true every time, one test still came out below 0.05, while none passed the Bonferroni threshold 0.05/20 = 0.0025. A/B testing tools that compare many metrics or many variants need this kind of correction.

Where you use the test checklist

  • Interviews: "which test would you use?" questions are answered with the four questions above, then the assumptions you would check.
  • A/B tests: a t-test (or Mann-Whitney U) on time spent and a chi-square or proportion test on conversions, with an effect size and a multiple-testing correction when many metrics are compared.
  • Data analysis reports: name the test, its assumptions, the statistic, the p-value and the effect size.
Watch out. A test is valid only when its assumptions hold. Check them before reading the p-value: independent observations; roughly normal data or a large sample for z, t and ANOVA (see Normality tests (Q-Q plot and Shapiro-Wilk)); expected counts of at least 5 for chi-square; and decide α, the tails and the test before looking at the data.
Try it yourself
  • Run stats.wilcoxon on the paired weights from Two-sample and paired t-tests and compare its p-value with the paired t-test's 0.5732.
  • In the large-sample example change size=200_000 to size=200. Is the 0.3-point difference still significant?
  • Change m = 20 to m = 100 in the multiple-testing example. How many false alarms appear at 0.05?

You understood something today that you didn't yesterday.