StatisticsSciPy 1.18 · pandas 3.0 · statsmodels 0.15 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
57 small wins to finish your pathNext lesson →

Confidence intervals

A confidence interval is a range computed from a sample as point estimate ± margin of error, by a procedure that captures the true population parameter in a stated share of repeated samples (the confidence level).

Last updated: 07 Oct, 2026 · SciPy 1.18

The Point estimates and standard error lesson ended with x̄ and its standard error. A confidence interval turns the two into a range of plausible values for μ, with a stated level of confidence.

Setting up the CAT score question

The video's question: on the quant test of the CAT exam, the population standard deviation is known to be 100. A sample of 25 test takers has a mean score of 520. Construct a 95% confidence interval for the mean.

  • Given: σ = 100, n = 25, x̄ = 520.
  • Confidence level 95%, so α = 1 − 0.95 = 0.05. The confidence level is always 1 − α.
  • Two tails: the interval leaves α/2 = 0.025 outside each end.

Because σ is known, the margin of error uses the z critical value zα/2 and the standard error σ/√n:

The z-interval for a mean, σ known
Building the 95% z-interval for the CAT scores · from the Complete Statistics for Data Science in 6 Hours video · 3:35:05 to 3:38:41

Building a 95% z-interval

The interval splits into an upper and a lower limit. For z0.025, the z-table is read at a left area of 1 − 0.025 = 0.975: row 1.9, column 0.06, so z0.025 = 1.96. The standard error is 100/√25 = 20.

The 95% z-interval for the CAT mean

So the 95% confidence interval for the mean quant score is 480.8 to 559.2, centred on the point estimate 520. The z-interval needs σ to be known, and either normally distributed scores or a sample large enough for the central limit theorem; with normal scores, n = 25 is fine. The Z-table and normal probabilities lesson covers reading the table.

ExampleRun on matplotlib 3.11.2
import numpy as np
import matplotlib.pyplot as plt
from scipy.stats import norm

x = np.linspace(440, 600, 400)
y = norm.pdf(x, loc=520, scale=20)        # x̄ ± its standard error of 20
lo, hi = norm.interval(0.95, loc=520, scale=20)
plt.figure(figsize=(8, 4))
plt.plot(x, y, color="black")
plt.fill_between(x, y, where=(x >= lo) & (x <= hi), alpha=0.3, label="central 95%")
for v in (lo, 520, hi):
    plt.axvline(v, color="red" if v != 520 else "grey", linestyle="--")
plt.title("95% z-interval for the CAT mean: 520 ± 1.96 × 20")
plt.xlabel("score")
plt.ylabel("density")
plt.legend()
plt.show()
print("limits:", round(lo, 1), "and", round(hi, 1))
A bell curve centred at 520 with standard deviation 20, the central 95% shaded between red dashed lines at 480.8 and 559.2.

The curve is the sampling distribution of x̄ with SE 20, drawn around the estimate 520. It describes sample means, not single CAT scores, which spread five times wider.

The t-interval when only the sample SD is known · from the Complete Statistics for Data Science in 6 Hours video · 3:42:25 to 3:45:27

Building a 95% t-interval when σ is unknown

The video's second question changes one fact: the same 25 test takers have mean 520 and a sample standard deviation s = 80, and the population SD is not given. With σ unknown and estimated by s, the critical value comes from the t distribution with n − 1 degrees of freedom:

The t-interval for a mean, σ unknown

Degrees of freedom: n − 1 = 25 − 1 = 24, the same n − 1 as in the sample variance. The t-table at df 24 and a one-tail area of 0.025 (two tails 0.05) gives 2.064. The standard error is 80/√25 = 16.

The 95% t-interval for the CAT mean

To two decimals the interval is 486.98 to 553.02. The t critical value 2.064 is larger than 1.96 because s is itself an estimate; the extra width pays for that uncertainty, and t approaches z as n grows. The One-sample t-test and the t distribution lesson covers the t distribution.

Computing both intervals in Python

norm.ppf(0.975) and t.ppf(0.975, df=24) replace the two table look-ups. SciPy can also return the limits in one call: norm.interval(0.95, loc=520, scale=20) and t.interval(0.95, df=24, loc=520, scale=16).

ExampleFrom the video, run on SciPy 1.18.1
import numpy as np
from scipy.stats import norm, t

xbar, n = 520, 25

# z-interval: the population SD sigma = 100 is known
z = norm.ppf(0.975)                  # area 0.975 to the left: the 1.96 of the z-table
se = 100 / np.sqrt(n)
print(f"z = {z:.4f}, SE = {se:.0f}, 95% z-interval: {xbar - z * se:.2f} to {xbar + z * se:.2f}")

# t-interval: only the sample SD s = 80 is known, df = n - 1 = 24
tc = t.ppf(0.975, df=n - 1)
se = 80 / np.sqrt(n)
print(f"t = {tc:.4f}, SE = {se:.0f}, 95% t-interval: {xbar - tc * se:.2f} to {xbar + tc * se:.2f}")
print(f"with the table's t = 2.064: {xbar - 2.064 * se:.3f} to {xbar + 2.064 * se:.3f}")
  • The z-interval is 480.80 to 559.20, as on the board, with z = 1.9600 and SE 20.
  • The t-interval is 486.98 to 553.02 with the exact t = 2.0639; the table's 2.064 gives 486.976 to 553.024.
  • The t-interval is narrower here only because s = 80 is smaller than σ = 100. With the same SD, the t-interval is always the wider one.

Reading a 95% confidence level correctly

A 95% confidence interval comes from a procedure that, over repeated samples, produces intervals containing the true μ 95% of the time. The 95% describes the method. Any one computed interval, such as 480.8 to 559.2, either contains μ or it does not; it is not true that μ lies inside it with probability 95%. The interval also says nothing about where 95% of individual scores fall.

A simulation makes this visible. The code below fixes the true mean at μ = 500 (σ = 100), draws 1,000 samples of 25, builds a 95% t-interval from each, and counts how many intervals contain 500.

ExampleRun on SciPy 1.18.1 and NumPy 2.5.3
import numpy as np
import matplotlib.pyplot as plt
from scipy import stats

rng = np.random.default_rng(42)
mu, sigma, n = 500, 100, 25          # the true mean is known here, so hits can be counted
intervals = []
for i in range(1000):
    s = rng.normal(mu, sigma, n)
    intervals.append(stats.t.interval(0.95, df=n - 1, loc=s.mean(), scale=stats.sem(s)))
hits = [lo <= mu <= hi for lo, hi in intervals]
print("intervals that contain mu = 500:", sum(hits), "of 1000")
print("misses among the first 50 drawn:", 50 - sum(hits[:50]))

plt.figure(figsize=(8, 5))
for i, (lo, hi) in enumerate(intervals[:50]):
    plt.plot([lo, hi], [i, i], color="tab:blue" if hits[i] else "red", linewidth=2)
plt.axvline(mu, color="black", linestyle="--")
plt.title("50 samples of 25, one 95% t-interval each (red misses mu = 500)")
plt.xlabel("score")
plt.ylabel("sample")
plt.show()
Fifty horizontal 95 percent t-intervals stacked from bottom to top, almost all blue and crossing a dashed vertical line at the true mean 500; the few red ones miss it.

What the coverage simulation shows

  • 948 of the 1,000 intervals contain μ: a coverage of 0.948, close to the promised 0.95. A longer run settles nearer to 0.95.
  • Each interval is different, because each sample has its own x̄ and s. Some are wider, some narrower.
  • 3 of the first 50 intervals miss (drawn in red). Nothing in a single interval says whether it is one of the misses, which is why the 95% belongs to the method and not to one interval.

Changing the width of a confidence interval

Three things set the width 2 × critical value × SE: the confidence level (a larger level needs a larger critical value), the spread σ, and the sample size n. The last line answers a planning question: how many test takers give a margin of error of ±10 points at 95%? Solving E = z σ/√n for n gives n = (z σ / E)².

ExampleRun on SciPy 1.18.1 and NumPy 2.5.3
import math
from scipy.stats import norm

sigma = 100
for conf in (0.90, 0.95, 0.99):
    z = norm.ppf(1 - (1 - conf) / 2)
    print(f"{conf:.0%}: z = {z:.3f}, width with n = 25: {2 * z * sigma / math.sqrt(25):.1f}, "
          f"with n = 100: {2 * z * sigma / math.sqrt(100):.1f}")

E = 10                                   # the margin of error wanted: ±10 points
n_needed = (norm.ppf(0.975) * sigma / E) ** 2
print("n for a ±10 margin at 95%:", round(n_needed, 2), "-> round up to", math.ceil(n_needed))
  • Higher confidence gives a wider interval: with n = 25 the width grows from 65.8 (90%) to 78.4 (95%) and 103.0 (99%).
  • Four times the data halves the width: 78.4 becomes 39.2 at 95%.
  • A ±10 margin needs n = 385: (1.96 × 100 / 10)² = 384.15, rounded up. This is how the video's interview question about the average size of sharks is answered in practice: decide the margin, use a pilot estimate of σ, and compute the sample size before collecting data.

z-interval vs t-interval

z-intervalt-interval
Whenσ known (normal data or large n)σ unknown, estimated by s
Formulax̄ ± zα/2 σ/√nx̄ ± tα/2, n−1 s/√n
Critical value at 95%1.962.064 for df 24, closer to 1.96 as n grows
CAT example480.8 to 559.2486.98 to 553.02
SciPynorm.interval(0.95, loc, scale)t.interval(0.95, df, loc, scale)

Where you use confidence intervals

  • Reporting an estimate: an average order value of ₹1,250 with a 95% CI of ₹1,180 to ₹1,320 says how precise the number is.
  • A/B tests: a CI for the difference between two conversion rates shows the size of the effect, not only whether there is one.
  • Planning a study: the sample size for a target margin of error, n = (z σ / E)².
Watch out. "There is a 95% probability that μ is between 480.8 and 559.2" is the most common misreading. μ is a fixed number; the randomness is in the sample. Say instead: "this interval came from a method that captures μ in 95% of samples".
Try it yourself
  • Change the confidence level to 99% in the CAT example (0.995 in both ppf calls). How much wider does each interval get?
  • In the coverage simulation, use stats.t.interval(0.90, ...). Does the count drop to about 900?
  • Set E = 5 in the width example. Why does halving the margin of error need four times the sample?

Every expert started right here.