StatisticsSciPy 1.18 · pandas 3.0 · statsmodels 0.15 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
57 small wins to finish your pathNext lesson →

Covariance

Covariance is a measure of how two numeric variables vary together: positive when they tend to rise and fall together, negative when one tends to fall as the other rises.

Last updated: 07 Oct, 2026 · SciPy 1.18

The tests so far compared groups. Often the question is about two measurements on the same things: do taller people weigh more, do students who study longer play less? Covariance is the first number that answers it.

Covariance and its formula · from the Complete Statistics for Data Science in 6 Hours video · 4:26:58 to 4:30:44

Spotting how two variables move together

The video's first table pairs weight X with height Y for four people. As weight goes up, height goes up; as weight goes down, height goes down. Its second table pairs hours of study with hours of play: as study goes up, play goes down.

Weight X (kg)Height Y (cm)Hours of study XHours of play Y
5016026
6017034
7018043
75181

Seeing the direction is easy with four rows. Covariance turns it into a number.

Computing covariance

For each pair, multiply how far x is from its mean by how far y is from its mean. The product is positive when both are above their means or both below, and negative when one is above and the other below. Covariance is the average of these products: divided by n for a whole population, and by n − 1 for a sample, as the board writes it, for the same reason as the Sample variance and why n − 1.

Sample covariance (divide by N for a population)

For the weight and height table, x̄ = 63.75 and ȳ = 172.75. The four products are 175.31, 10.31, 45.31 and 92.81, all positive, and they add up to 323.75:

The weight and height table, as a sample

Divided by n = 4 instead, the population covariance is 80.94. The study and play table gives −1.5: a negative covariance for a falling relationship. The notes that go with the video add a third small example, X = 2, 4, 6 and Y = 3, 5, 7, whose covariance is ((−2)(−2) + 0 + (2)(2))/2 = 4.

ExampleFrom the video's board, run on matplotlib 3.11.2
import numpy as np
import matplotlib.pyplot as plt

weight = np.array([50, 60, 70, 75])
height = np.array([160, 170, 180, 181])
sign = (weight - weight.mean()) * (height - height.mean())
plt.figure(figsize=(6, 4.5))
plt.axvline(weight.mean(), color="grey", linestyle="--")
plt.axhline(height.mean(), color="grey", linestyle="--")
plt.scatter(weight, height, s=80, c=np.where(sign > 0, "green", "red"))
for x, y, s in zip(weight, height, sign):
    plt.annotate(f"{s:+.1f}", (x, y), textcoords="offset points", xytext=(8, -12))
plt.xlim(45, 80)
plt.ylim(155, 185)
plt.title("Each point's (x - x̄)(y - ȳ): all four are positive")
plt.xlabel("weight (kg), dashed line at x̄ = 63.75")
plt.ylabel("height (cm), dashed line at ȳ = 172.75")
plt.show()
print("sum of products:", sign.sum(), " / (n - 1) =", round(sign.sum() / 3, 4))
The four weight and height points with dashed lines at the mean weight 63.75 and mean height 172.75; every point is in the upper-right or lower-left quarter, all four are green, and each is labelled with its positive product: 175.3, 10.3, 45.3 and 92.8.

Reading the sign of covariance

  • Positive: x and y tend to rise together and fall together (weight and height).
  • Negative: when x rises, y tends to fall (study and play).
  • Zero: no linear relationship. That is not the same as no relationship: points on a U-shape have a covariance near 0, because the products on the left and right halves cancel, even though y depends on x completely.
ExampleRun on matplotlib 3.11.2
import numpy as np
import matplotlib.pyplot as plt

rng = np.random.default_rng(42)
x = rng.uniform(-3, 3, 200)
panels = {"rising": x + rng.normal(0, 0.8, 200),
          "falling": -x + rng.normal(0, 0.8, 200),
          "no pattern": rng.normal(0, 1.5, 200),
          "U-shape": x ** 2 + rng.normal(0, 0.5, 200)}
fig, axes = plt.subplots(1, 4, figsize=(13, 3.2))
for ax, (name, y) in zip(axes, panels.items()):
    c = np.cov(x, y)[0, 1]
    ax.scatter(x, y, s=8)
    ax.set_title(f"{name}: cov = {c:.2f}")
    ax.set_xlabel("x")
axes[0].set_ylabel("y")
plt.tight_layout()
plt.show()
print({name: round(float(np.cov(x, y)[0, 1]), 2) for name, y in panels.items()})
Four scatter plots of 200 points: a rising cloud with positive covariance, a falling cloud with negative covariance, a shapeless cloud with covariance near 0, and a U-shape with covariance near 0 although y clearly depends on x.

Computing covariance in Python

Deviation products by hand

python
import numpy as np

weight = np.array([50, 60, 70, 75])          # X
height = np.array([160, 170, 180, 181])      # Y
products = (weight - weight.mean()) * (height - height.mean())

Dividing by n − 1 or by n

python
cov_sample = products.sum() / (len(weight) - 1)    # divide by n - 1
cov_population = products.sum() / len(weight)      # divide by n
ExampleFrom the video, run on NumPy 2.5.3
print("deviation products:", products)
print("sample covariance:", round(cov_sample, 4), " population:", round(cov_population, 4))
print("np.cov:", np.cov(weight, height))           # 2 x 2 matrix, n - 1 by default
study, play = [2, 3, 4], [6, 4, 3]
print("study vs play:", np.cov(study, play)[0, 1])
print("X = 2, 4, 6 and Y = 3, 5, 7:", np.cov([2, 4, 6], [3, 5, 7])[0, 1])
  • The sample covariance is 107.9167 and the population covariance 80.9375.
  • np.cov returns a 2 × 2 matrix. The off-diagonal entries are cov(weight, height) = 107.9167; the diagonal holds each variable's own sample variance, 122.9167 for weight and 96.9167 for height.
  • Study vs play gives −1.5 and the notes' example gives 4.0, as worked above.

Why the size of a covariance is hard to read

Covariance carries the units of x times the units of y: kg·cm here. Measure height in metres instead of centimetres and the covariance shrinks 100 times, though the relationship is the same. So, as the video says, a covariance of +1000 is not "more related" than +100, and two covariances can only be compared when the units match. pandas shows it, and also shows that the covariance of a variable with itself is its variance:

ExampleRun on pandas 3.0.6
import pandas as pd

df = pd.DataFrame({"weight_kg": [50, 60, 70, 75], "height_cm": [160, 170, 180, 181]})
df["height_m"] = df["height_cm"] / 100
print(df.cov().round(4))                     # pandas: ddof=1
print("var(weight) =", df["weight_kg"].var().round(4), "= cov(weight, weight)")
Covariance sits on an open number line with values such as -2000, -200, 0, +100 and +1000 whose size depends on the units; Pearson r = cov(X, Y)/(sigma X sigma Y) sits on a closed scale from -1, a perfect negative line, through 0, no linear relationship, to +1, a perfect positive line.

Dividing by the two standard deviations removes the units and fixes the range to −1 to +1. That is the Pearson correlation coefficient.

Covariance vs Pearson correlation

CovariancePearson correlation r
Formulaaverage of (x − x̄)(y − ȳ)cov(x, y) / (s_x s_y)
Rangeany value−1 to +1
Unitsunits of x × units of ynone
Tells youthe directionthe direction and the strength of a linear relationship
Weight and height table107.92 (n − 1)0.989

Where you use covariance

  • Covariance matrices: Principal component analysis (PCA) finds the directions of largest variance from the covariance matrix of the features.
  • Portfolio risk: the variance of a sum of returns adds twice the covariance of each pair, so assets with negative covariance reduce risk.
  • Sums of variables: Var(X + Y) = Var(X) + Var(Y) + 2 cov(X, Y), so Var(X + Y) = Var(X) + Var(Y) only when the covariance is 0.
Watch out. np.cov divides by n − 1 by default (bias=True gives n), while np.var divides by n by default (ddof=0). pandas .cov() and .var() both use n − 1. Mixing them gives a covariance matrix whose diagonal does not match your variances.
Try it yourself
  • Add a fifth person, weight 80 kg and height 150 cm, to the arrays. Does the covariance stay positive?
  • Run np.cov(weight, height, bias=True). Is the off-diagonal entry 80.9375?
  • Change height_m to height in millimetres (×10). How many times larger does its covariance with weight get?

Slow is fine. Stopping is the only problem.