Statistics interview questions
Statistics interview questions are the questions data science interviews use to check that a candidate can define the core ideas, choose the right method and compute and interpret a result, by hand and in Python.
Last updated: 07 Oct, 2026 · SciPy 1.18
The questions below come from the interview notes that go with the video, from one-line definitions to full A/B test use cases. Each answer links to the lesson that explains it.
Answering the core questions
- What is the central limit theorem and why does it matter? For independent, identically distributed values with a finite variance, the mean of a sample of size n is close to normal for large n, with mean μ and standard error σ/√n, whatever the population's shape. That is why z- and t-based tests and intervals work on non-normal data. See Central limit theorem.
- What are Type I and Type II errors? A Type I error rejects a true H₀ (a false positive); its probability is α. A Type II error fails to reject a false H₀ (a false negative); its probability is β, and the power of the test is 1 − β. See Type I and Type II errors.
- What is a p-value? The probability, if H₀ is true, of a result at least as extreme as the one observed. Reject H₀ when p ≤ α; otherwise fail to reject it. It is not the probability that H₀ is true. See P-value.
- What is R-squared? The share of the variance of the target that a regression model explains: 1 is a perfect fit, 0 explains nothing. See R squared and adjusted R squared.
- What is the difference between correlation and causation? Correlation means two variables move together; causation means changing one changes the other. Ice cream sales and drownings are correlated because hot weather drives both. See Pearson correlation coefficient.
- Parametric or non-parametric? Parametric tests assume a form for the population, usually normal, and test its parameters (t-tests, ANOVA). Non-parametric tests do not (Mann-Whitney U, Kruskal-Wallis, chi-square on counts). See Choosing a statistical test.
- Cross-validation or bootstrapping? Cross-validation splits the data into folds to estimate how a model does on new data. Bootstrapping resamples the data with replacement to estimate the spread of a statistic. See Cross-validation.
Answering descriptive statistics questions
- Mean, median and mode: the average, the middle value of the sorted data and the most frequent value. For skewed data or outliers the median is the better centre. See Mean, median and mode.
- Population and sample: the population is the whole group; a sample is a subset. Numbers that describe a population are parameters, numbers computed from a sample are statistics. See Population and sample.
- Variance and standard deviation: variance is the average squared distance from the mean (divided by n − 1 for a sample); the standard deviation is its square root, in the data's own units. See Variance and standard deviation.
- IQR: Q3 − Q1, the spread of the middle half; the outlier fences sit 1.5 × IQR beyond the quartiles. See Quartiles and the interquartile range.
- Box plot: the box runs from Q1 to Q3 with the median inside. The whiskers reach the most extreme values within 1.5 × IQR of the box, and points beyond are drawn as outliers. See Five-number summary and box plot.
- Skewness and kurtosis: skewness measures asymmetry (a long right tail is positive skew); kurtosis measures how heavy the tails are compared with a normal curve. See Skewness and kurtosis.
- z-score: (x − μ)/σ, how many standard deviations a value is from the mean. See Z-score and the standard normal distribution.
- Covariance and correlation: covariance gives the direction in units of x × y; Pearson's r divides by both standard deviations to give a unit-free strength from −1 to +1. See Covariance.
- Simpson's paradox: a trend that appears in every group can reverse when the groups are combined, because a third variable is unevenly spread across them.
- Why the standard deviation can mislead: one outlier inflates it; the IQR or the median absolute deviation is more resistant. See Range, MAD and coefficient of variation.
Working the descriptive use cases
Which region has the highest and the steadiest sales?
Monthly sales in thousands for three regions over seven months. The highest mean answers the first part, the smallest standard deviation the second.
import pandas as pd
sales = pd.DataFrame({"North": [12, 15, 14, 13, 17, 19, 20],
"South": [22, 21, 20, 23, 25, 26, 28],
"West": [32, 30, 31, 29, 30, 33, 35]})
print(pd.DataFrame({"mean": sales.mean(),
"SD (n)": sales.std(ddof=0),
"SD (n - 1)": sales.std()}).round(3))mean SD (n) SD (n - 1) North 15.714 2.814 3.039 South 23.571 2.665 2.878 West 31.429 1.917 2.070
West has the highest mean, 31.429, and the smallest spread on either convention: 1.917 dividing by n or 2.070 dividing by n − 1. North's SD is 2.814 (n) and South's 2.665 (n). Name the divisor you use.
Which recovery group is quickest and least variable?
import numpy as np
groups = {"A": [5, 6, 4, 5, 7, 5, 6], "B": [7, 8, 7, 9, 8, 7, 9], "C": [5, 7, 6, 5, 6, 6, 5]}
for name, v in groups.items():
q1, q3 = np.percentile(v, [25, 75]) # linear method (default)
w1, w3 = np.percentile(v, [25, 75], method="weibull") # the (n + 1)p hand method
print(f"{name}: median {np.median(v)}, range {max(v) - min(v)}, "
f"IQR linear {q3 - q1}, IQR (n + 1)p {w3 - w1}")A: median 5.0, range 3, IQR linear 1.0, IQR (n + 1)p 1.0 B: median 8.0, range 2, IQR linear 1.5, IQR (n + 1)p 2.0 C: median 6.0, range 2, IQR linear 1.0, IQR (n + 1)p 1.0
Group A has the quickest median recovery, 5 days. B and C share the smallest range, 2 days. The IQR depends on the quartile method: for group B the (n + 1)p method gives Q1 = 7, Q3 = 9 and an IQR of 2, while NumPy's default linear method gives 1.5. Say which method you used (see Quartiles and the interquartile range).
How much does one outlier move the mean?
Day 1 page-load times are 3, 2.5, 2.8, 3.1 and 15 seconds, the last from a server glitch. With it the mean is 26.4/5 = 5.28 s; without it, 11.4/4 = 2.85 s, a jump of 2.43 s from one value. The median, 3 s with the outlier, barely moves. A log transform pulls the long tail in: ln 3 = 1.0986, and the glitch's ln 15 = 2.708 sits much closer to it than 15 does to 3.
Working the inferential use cases
A z-test on exam scores
An exam board says students in state X average 52 in mathematics. A sample of 100 students averages 54 with a standard deviation of 10. At α = 0.05, is the average different from 52? H₀: μ = 52, H₁: μ ≠ 52, two-tailed.
import numpy as np
from scipy import stats
xbar, mu0, s, n = 54, 52, 10, 100
z = (xbar - mu0) / (s / np.sqrt(n))
print("z =", z)
print("two-tailed p =", round(2 * stats.norm.sf(abs(z)), 4))
print("one-tailed area =", round(stats.norm.sf(abs(z)), 4), "(not the answer to a 'differs' question)")
print("t-test p, df 99 =", round(2 * stats.t.sf(abs(z), n - 1), 4))z = 2.0 two-tailed p = 0.0455 one-tailed area = 0.0228 (not the answer to a 'differs' question) t-test p, df 99 = 0.0482
z = (54 − 52)/(10/√100) = 2.0 and the two-tailed p-value is 0.0455 ≤ 0.05, so we reject H₀: the state's average differs from 52. The area in one tail, 0.0228, answers a one-sided question, not this one. Strictly, 10 is a sample SD, so this is a t-test with 99 degrees of freedom (p = 0.0482); with n = 100 the two agree closely. See One-sample z-test.
Paired, ANOVA and chi-square questions
- An energy drink tested on 15 people before and after: the same people are measured twice, so it is a paired t-test,
stats.ttest_rel(after, before). See Two-sample and paired t-tests. - Three fertilizers and crop yield: three groups of a numeric variable, so one-way ANOVA,
stats.f_oneway(a, b, c), then Tukey's HSD if it rejects. See One-way ANOVA (F-test). - Gender and product preference for 100 customers, [[30, 20], [25, 25]]: a chi-square test of independence.
chi2_contingencygives χ² = 0.646 and p = 0.421 with Yates' correction (1.010 and 0.315 without), so we fail to reject H₀: no evidence that preference depends on gender. See Chi-square test of independence. - The Standard Chartered ATM question from the video: set H₀ as "mean daily withdrawals are at most the break-even amount" against H₁ "more", collect a sample of days and run a one-tailed one-sample t-test. See One-sample t-test and the t distribution.
An A/B test on time spent and purchases
An online store shows 8 users the old design (A) and 8 the new design (B). It records minutes on the page and whether each user bought. Two questions: does time spent differ, and is buying related to the design?
import numpy as np
from scipy import stats
time_a = [3, 5, 4, 6, 5, 5, 6, 4] # old design, minutes
time_b = [6, 7, 7, 7, 8, 6, 7, 8] # new design
print("means:", np.mean(time_a), np.mean(time_b),
" SDs:", round(np.std(time_a, ddof=1), 3), round(np.std(time_b, ddof=1), 3))
r = stats.ttest_ind(time_a, time_b, equal_var=False)
print(f"Welch t = {r.statistic:.3f}, df = {r.df:.1f}, p = {r.pvalue:.5f}")
purchases = np.array([[4, 4], # group A: no purchase, purchase
[1, 7]]) # group B
c = stats.chi2_contingency(purchases, correction=False)
print("expected:", c.expected_freq.tolist())
print(f"chi-square = {c.statistic:.3f}, p = {c.pvalue:.3f}")
print(f"Fisher's exact p = {stats.fisher_exact(purchases).pvalue:.3f}")means: 4.75 7.0 SDs: 1.035 0.756 Welch t = -4.965, df = 12.8, p = 0.00027 expected: [[2.5, 5.5], [2.5, 5.5]] chi-square = 2.618, p = 0.106 Fisher's exact p = 0.282
- Time spent: means 4.75 and 7.0 minutes, SDs 1.035 and 0.756. Welch's t = −4.965 with p = 0.00027, so we reject H₀: users spend longer on the new design. With equal group sizes the pooled test gives the same t, with df = 14 and critical value ±2.145.
- Purchases: the table [[4, 4], [1, 7]] gives χ² = 2.618 and p = 0.106 without the correction, so we fail to reject H₀.
- Expected counts of 2.5 break the chi-square rule of at least 5. Fisher's exact test is the right tool for this table, and its p = 0.282 also fails to reject. Eight users per group is too few to judge purchases.
Recalling formulas vs deriving them
| Question | Formula to recall | Lesson |
|---|---|---|
| z-test for a mean | z = (x̄ − μ₀)/(σ/√n) | One-sample z-test |
| t-test for a mean | t = (x̄ − μ₀)/(s/√n), df n − 1 | One-sample t-test and the t distribution |
| Proportion | z = (p̂ − p₀)/√(p₀(1 − p₀)/n) | Z-test for a proportion |
| Chi-square | Σ(O − E)²/E, df k − 1 or (r − 1)(c − 1) | Chi-square goodness-of-fit test |
| Correlation | r = cov(x, y)/(s_x s_y) | Pearson correlation coefficient |
Where you use these answers
- Screening rounds: short definitions (p-value, Type I error, CLT) in one or two exact sentences.
- Case rounds: a use case such as the A/B test, answered with hypotheses, the test and why, the assumptions, the numbers and a plain decision.
- Take-home tasks: the same steps in Python, with the library defaults named (ddof, Yates, equal_var).
Related
- Previous: Spearman rank correlation
- Start of the course: Descriptive and inferential statistics
- Add an eighth month of 40 to the West region. Is it still the steadiest region?
- In the exam question change the sample mean to 53.5. Does the two-tailed test still reject H₀?
- Double the A/B purchase table to [[8, 8], [2, 14]] and run the chi-square and Fisher tests again.
Little by little, you're building something great.