StatisticsSciPy 1.18 · pandas 3.0 · statsmodels 0.15 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
57 small wins to finish your pathNext lesson →

Chi-square test of independence

A chi-square test of independence is a hypothesis test that checks whether two categorical variables are related, by comparing the counts in their contingency table with the counts expected if they were independent.

Last updated: 07 Oct, 2026 · SciPy 1.18

The Chi-square goodness-of-fit test tested one variable against stated shares. The video's notebook asks a question about two variables at once: in a restaurant's tips data, is being a smoker related to the diner's sex?

Building the contingency table

A contingency table (a crosstab) counts every combination of the two variables: rows for one variable, columns for the other. The notebook builds it from seaborn's tips dataset, 244 restaurant bills.

python
import pandas as pd
import seaborn as sns
from scipy import stats

dataset = sns.load_dataset("tips")
dataset_table = pd.crosstab(dataset["sex"], dataset["smoker"])
print(dataset_table)

The table has Male: 60 smokers and 97 non-smokers, Female: 33 and 54. The hypotheses are H₀: sex and smoking are independent, and H₁: they are related.

Computing expected counts under independence

If the two variables were independent, the share of smokers would be the same for men and women: the overall share, 93 of 244. So the expected count for a cell is its row total times its column total divided by the grand total N. For Male and smoker: 157 × 93 / 244 = 59.84.

Expected counts, the statistic and the degrees of freedom for an r × c table
Smoker: YesSmoker: NoRow total
Male, observed (expected)60 (59.84)97 (97.16)157
Female, observed (expected)33 (33.16)54 (53.84)87
Column total93151244

Every observed count is within 0.16 of its expected count. A 2 × 2 table has (2 − 1)(2 − 1) = 1 degree of freedom: once one cell is fixed, the totals fix the other three.

Running the test by hand

The notebook computes the statistic, the critical value and the p-value itself. Here is the same calculation with the expected counts written out from the formula above.

Observed counts, expected counts and the statistic

python
observed = dataset_table.values
row_totals = observed.sum(axis=1, keepdims=True)
col_totals = observed.sum(axis=0, keepdims=True)
expected = row_totals * col_totals / observed.sum()   # row total x column total / N
chi_square = ((observed - expected) ** 2 / expected).sum()
ddof = (observed.shape[0] - 1) * (observed.shape[1] - 1)
ExampleFrom the video's notebook, run on SciPy 1.18.1
alpha = 0.05
critical_value = stats.chi2.ppf(q=1 - alpha, df=ddof)
p_value = stats.chi2.sf(chi_square, df=ddof)
print("expected:\n", expected.round(2))
print(f"chi-square = {chi_square:.6f}, df = {ddof}, critical = {critical_value:.3f}, p = {p_value:.4f}")
print("reject H0" if p_value <= alpha else "fail to reject H0: no evidence of a relationship")

χ² = 0.001935 is far below the critical 3.841, and p = 0.9649. If sex and smoking were independent, a table at least this far from the expected counts would turn up in about 96% of samples of 244. We fail to reject H₀: the data show no relationship between sex and smoking.

Running chi2_contingency and the Yates correction

scipy.stats.chi2_contingency does the whole test in one call. On a 2 × 2 table it applies Yates' continuity correction by default (correction=True): it moves every observed count 0.5 closer to its expected count before squaring, to make the χ² approximation less eager to reject on small tables. SciPy never moves a count past its expected value, so when |O − E| is under 0.5 the corrected gap is 0.

ExampleFrom the video's notebook, run on SciPy 1.18.1
with_yates = stats.chi2_contingency(dataset_table)                   # correction=True by default
no_yates = stats.chi2_contingency(dataset_table, correction=False)
print("default (Yates):  ", round(with_yates.statistic, 6), round(with_yates.pvalue, 4), with_yates.dof)
print("correction=False: ", round(no_yates.statistic, 6), round(no_yates.pvalue, 4), no_yates.dof)
print("largest |O - E|:  ", round(abs(dataset_table.values - with_yates.expected_freq).max(), 3))
  • The default call returns χ² = 0.0 and p = 1.0. Every |O − E| here is 0.16, below 0.5, so the correction removes every gap.
  • correction=False returns 0.001935 and p = 0.9649, the hand calculation.
  • Both say fail to reject H₀. The correction matters on small tables near the decision line, so say which version you report.

Finding a relationship: day and meal time

The same data with two variables that are clearly related: the day of the week and whether the bill was for lunch or dinner.

ExampleRun on SciPy 1.18.1
import pandas as pd
import seaborn as sns
from scipy import stats

tips = sns.load_dataset("tips")
day_time = pd.crosstab(tips["day"], tips["time"])
print(day_time)
r = stats.chi2_contingency(day_time)
print(f"chi-square = {r.statistic:.2f}, df = {r.dof}, p = {r.pvalue:.3g}")
print("smallest expected count:", r.expected_freq.min().round(2))

χ² = 217.11 with df = (4 − 1)(2 − 1) = 3 and p = 8.45 × 10⁻⁴⁷, so we reject H₀: meal time depends on the day. Thursday is almost all lunch, Saturday and Sunday are all dinner. The smallest expected count is 5.30, so the test is valid. The test says the variables are related; the table shows how.

Comparing the smoker shares

ExampleRun on matplotlib 3.11.2
import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt

tips = sns.load_dataset("tips")
share = pd.crosstab(tips["sex"], tips["smoker"], normalize="index")["Yes"]
plt.figure(figsize=(6, 4))
plt.bar(share.index.astype(str), share.values, color=["tab:blue", "tab:orange"])
plt.axhline(93 / 244, color="grey", linestyle="--", label="all diners: 93 of 244")
plt.ylim(0, 0.6)
plt.title("Share of smokers by sex in the tips data")
plt.ylabel("share who smoke")
plt.legend()
plt.show()
print(share.round(3))
Two bars of almost equal height, the share of smokers among male diners (0.382) and among female diners (0.379), both on the dashed line for all diners, 93 of 244.

38.2% of men and 37.9% of women in the data smoke, both close to the overall 38.1%. That is what "independent" looks like, and why χ² is near 0.

Chi-square test of independence vs goodness of fit

Test of independenceGoodness of fit
QuestionAre two categorical variables related?Does one variable match stated shares?
Tabler × c contingency tableOne row of k counts
Expected countRow total × column total / Nn × stated share
Degrees of freedom(r − 1)(c − 1)k − 1
SciPychi2_contingency(table, correction=...)chisquare(f_obs, f_exp)

Where you use a chi-square test of independence

  • Survey analysis: is product preference related to gender, region or age band?
  • A/B tests on yes/no outcomes: is buying related to which page design a user saw?
  • Feature screening: is a categorical feature related to a categorical target before it goes into a model?
Watch out. Every expected count should be at least 5. On a small 2 × 2 table with an expected count below 5, use Fisher's exact test, stats.fisher_exact(table). A significant χ² also says nothing about cause, only that the variables are not independent in this data.
Try it yourself
  • Replace the tips table with [[30, 20], [25, 25]] (gender by product preference, 100 customers) and run chi2_contingency with and without the correction.
  • Cross tips["smoker"] with tips["time"]. Is smoking related to lunch or dinner?
  • Run stats.fisher_exact(dataset_table) on the sex by smoker table and compare its p-value with 0.9649.

Every expert started right here.