StatisticsSciPy 1.18 · pandas 3.0 · statsmodels 0.15 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
57 small wins to finish your pathNext lesson →

Spearman rank correlation

Spearman's rank correlation coefficient is a measure of the strength and direction of a monotonic relationship between two variables: the Pearson correlation computed on the ranks of the values instead of the values themselves.

Last updated: 07 Oct, 2026 · SciPy 1.18

The Pearson correlation coefficient asks how close the points are to a straight line. Many relationships are not straight but still always go one way: more experience, higher salary, but with diminishing steps. Spearman's correlation measures that.

Why Pearson falls short on a curve · from the Complete Statistics for Data Science in 6 Hours video · 4:36:07 to 4:37:20

Seeing where Pearson falls short

The video's figure shows points where y rises every time x rises, quickly at first and then by small amounts. Pearson gives 0.88 there, not 1, because the points do not lie on a line. A relationship like this, where y only ever rises (or only ever falls) as x rises, is called monotonic. It need not be linear. Spearman's correlation is 1 for any perfectly rising monotonic relationship and −1 for any perfectly falling one.

ExampleRun on SciPy 1.18.1
import numpy as np
import matplotlib.pyplot as plt
from scipy import stats

x = np.linspace(0, 1, 40)
y = np.exp(6 * x)                            # always rising, but not a straight line
r = stats.pearsonr(x, y).statistic
r_s = stats.spearmanr(x, y).statistic
plt.figure(figsize=(6, 4))
plt.scatter(x, y, s=15)
plt.title(f"A monotonic curve: Pearson r = {r:.2f}, Spearman = {r_s:.2f}")
plt.xlabel("x")
plt.ylabel("y = exp(6x)")
plt.show()
print(f"Pearson {r:.3f}, Spearman {r_s:.3f}")
Forty points on the curve y = exp(6x), flat at the left and rising steeply at the right; the title shows Pearson r well below 1 and Spearman exactly 1.

The curve never goes down, so the ranks of x and y agree perfectly and Spearman is 1.00, while Pearson is pulled below 1 by the bend.

Ranking the values

Ranking the height and weight table · from the Complete Statistics for Data Science in 6 Hours video · 4:38:45 to 4:40:29

Spearman replaces every value by its rank and then computes Pearson's formula on the ranks: the covariance of the ranks divided by the standard deviations of the ranks. The board's example adds a fifth person to the height and weight table and ranks each column with the highest value as rank 1.

The heights 170, 160, 150, 145, 180 get ranks 2, 3, 4, 5, 1 and the weights 75, 62, 60, 55, 85 get ranks 2, 3, 4, 5, 1; the rank columns match row for row, so the Spearman correlation is 1.
Spearman's correlation: Pearson's r on the ranks R(x) and R(y)

When no values are tied there is a shortcut using d, the difference between a row's two ranks:

The no-ties shortcut

Ranking from the highest or from the lowest gives the same r_s, as long as both columns are ranked the same way. In the board's table every d is 0, so r_s = 1 − 0 = 1.

Computing Spearman correlation in Python

Ranks with rankdata

python
import numpy as np
from scipy import stats

height = np.array([170, 160, 150, 145, 180])
weight = np.array([75, 62, 60, 55, 85])
rank_h = stats.rankdata(-height)        # highest height gets rank 1, as on the board
rank_w = stats.rankdata(-weight)
ExampleFrom the video, run on SciPy 1.18.1
print("R(height):", rank_h, " R(weight):", rank_w)
print("Pearson on the ranks:", round(stats.pearsonr(rank_h, rank_w).statistic, 5))
print("spearmanr:", round(stats.spearmanr(height, weight).statistic, 5))
print("Pearson on the raw values:", round(stats.pearsonr(height, weight).statistic, 5))
  • Both rank columns print as 2, 3, 4, 5, 1, the board's ranks.
  • Pearson on the ranks and spearmanr both give 1.0. The order of the heights matches the order of the weights exactly.
  • Pearson on the raw values is 0.97663: high, but below 1, because the points are not on one straight line.

The marks example with the shortcut

The notes that go with the video rank a second table, X = 1, 3, 7, 0, 8 and Y = 2, 4, 5, 7, 1, with ranks R_x = 4, 3, 2, 5, 1 and R_y = 4, 3, 2, 1, 5. The last two rows swap places, so d = 0, 0, 0, 4, −4 and Σd² = 32:

The marks example
ExampleFrom the video's notes, run on SciPy 1.18.1
import numpy as np
from scipy import stats

X = np.array([1, 3, 7, 0, 8])
Y = np.array([2, 4, 5, 7, 1])
d = stats.rankdata(-X) - stats.rankdata(-Y)       # rank differences
n = len(X)
r_s = 1 - 6 * (d ** 2).sum() / (n * (n ** 2 - 1))
print("ranks X:", stats.rankdata(-X), " ranks Y:", stats.rankdata(-Y))
print("sum of d^2:", (d ** 2).sum(), " shortcut r_s:", round(r_s, 4))
print("spearmanr:", round(stats.spearmanr(X, Y).statistic, 4),
      " Pearson:", round(stats.pearsonr(X, Y).statistic, 4))

Ties get average ranks

When two values are equal they share the average of the ranks they would take. The shortcut formula is then only approximate, while spearmanr stays exact because it computes Pearson on the average ranks.

ExampleRun on SciPy 1.18.1
from scipy import stats

ratings = [5, 3, 3, 4, 1]                    # two people gave 3
print("average ranks:", stats.rankdata(ratings))

Comparing Spearman and Pearson on outliers and curves

ExampleRun on SciPy 1.18.1
import numpy as np
from scipy import stats

rng = np.random.default_rng(42)
x = np.arange(1, 21)
y = x + rng.normal(0, 2, 20)
y_out = y.copy()
y_out[-1] = -60                              # one wrong entry
u = np.linspace(-2, 2, 21)
for name, a, b in (("clean", x, y), ("one outlier", x, y_out), ("U-shape", u, u ** 2)):
    print(f"{name:12} Pearson {stats.pearsonr(a, b).statistic:6.3f}   "
          f"Spearman {stats.spearmanr(a, b).statistic:6.3f}")

On clean, roughly linear data the two agree. One wrong entry drags Pearson far down, while Spearman, which only sees that the point has the lowest rank, moves much less. On a U-shape both are near 0: Spearman captures curves that keep one direction, not every non-linear relationship.

Spearman vs Pearson correlation

Spearman r_sPearson r
Measuresmonotonic associationlinear association
Computed onthe ranksthe values
Outlierslittle effectcan change r a lot
Datanumeric or ordinalnumeric
Board's height and weight1.00.977
Pythonspearmanr, df.corr(method="spearman")pearsonr, df.corr()

Where you use Spearman rank correlation

  • Ordinal data: ratings, survey answers on a 1 to 5 scale, or rankings from two judges.
  • Skewed data with outliers: income, house prices or page-load times, where a few large values dominate Pearson.
  • Monotonic but curved relationships: dose and response, experience and salary.
Watch out. Spearman captures monotonic relationships, not every non-linear one: a U-shape or a wave gives r_s near 0 even when y depends on x. Plot the data before trusting either coefficient.
Try it yourself
  • In the board's table change the 180 cm person's weight to 50 kg. What happens to Spearman and to Pearson?
  • Run stats.kendalltau(height, weight), another rank correlation, on the board's table.
  • In the curve example replace np.exp(6 * x) with np.sin(6 * x). Why does Spearman drop below 1?

Every expert started right here.