Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Silhouette score

The silhouette score is a clustering metric that rates how well each point fits its own cluster compared with the nearest other cluster, on a scale from −1 to +1.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

A classification model is checked against the true labels with a confusion matrix, accuracy, precision or recall. A clustering model has no true labels, so the video validates it with the silhouette score. It works for K-means clustering and for Hierarchical clustering, and it checks the K that the elbow method picked.

Measuring a(i) inside the cluster

Silhouette: measuring a(i) · from the Complete Machine Learning in 6 Hours video · 311:34 to 313:28

Take one point i of cluster C1. a(i) is the average distance from i to every other point of C1. The sum is divided by |C1| − 1 because i itself is left out. In this clip i is called the centroid; i is one data point of the cluster, as the video says at the end of the next clip.

Measuring b(i) and the score s(i)

Silhouette: b(i) and the score from −1 to +1 · from the Complete Machine Learning in 6 Hours video · 313:28 to 317:01

b(i) looks outside the cluster. For every other cluster, average the distance from i to all of its points, and keep the smallest of those averages: the nearest other cluster.

In a good clustering a point is close to its own cluster and far from the next one, so b(i) is much larger than a(i). The board first writes a(i) ≫ b(i) as the good case and then crosses it out; the good case is b(i) ≫ a(i).

The board notes also write s(i) in three pieces, which shows the sign at a glance:

When a(i) is much smaller than b(i), the first case is close to 1. When a(i) is much larger, the last case is close to −1. So −1 ≤ s(i) ≤ 1 always holds, and the nearer the score is to +1, the better the clustering model.

s(i) runs from −1 to +1. Near +1 the point sits well inside its cluster. Near 0 it lies on the border between two clusters, and the clustering still needs improving. Below 0, a(i) is larger than b(i): the point is closer to another cluster than to its own. The silhouette score of a whole model is the mean s(i) over all points.

Three small clusters: point i in C1 with short red lines to its own cluster, a(i) = 1.00, dashed lines to the nearest cluster C2, b(i) = 4.37, and a scale from minus one to plus one with s(i) = 0.77 near the good end.

Computing s(i) by hand and with silhouette_samples

The diagram's nine points, three per cluster, make the formulas small enough to follow. Point i = (2, 2) is 1 away from both of its neighbours in C1, so a(i) = 1. Its distances to C2 are 4, 5 and √17 ≈ 4.12, an average of 4.37; C3 is farther, at 6.36 on average, so b(i) = 4.37 and s(i) = (4.37 − 1) / 4.37 = 0.77. The code repeats the sum and compares it with scikit-learn.

ExampleThe board's formulas on nine points, run on scikit-learn 1.9.1
import numpy as np
from sklearn.metrics import silhouette_samples, silhouette_score

X = np.array([[2, 2], [1, 2], [2, 1],      # C1, point i is the first one
              [6, 2], [7, 2], [6, 1],      # C2
              [2, 8], [3, 8], [2, 9]])     # C3
labels = np.array([0, 0, 0, 1, 1, 1, 2, 2, 2])

i = X[0]
d = np.sqrt(((X - i) ** 2).sum(axis=1))     # distance from i to every point
a = d[1:3].mean()                            # own cluster, without i itself
b = min(d[labels == 1].mean(), d[labels == 2].mean())   # nearest other cluster
s = (b - a) / max(a, b)
print("a(i) =", round(a, 3), " b(i) =", round(b, 3), " s(i) =", round(s, 3))
print("silhouette_samples:", round(silhouette_samples(X, labels)[0], 3))
print("silhouette_score  :", round(silhouette_score(X, labels), 3))
  • The hand result and silhouette_samples agree at 0.771 for point i.
  • silhouette_score is 0.782, the mean s(i) over all nine points: every cluster is tight and far from the others.

Silhouette analysis on the make_blobs data

The silhouette analysis code · from the Complete Machine Learning in 6 Hours video · 329:19 to 331:30

The practical keeps the make_blobs data from K-means clustering and loops over 2 to 6 clusters with code taken from the scikit-learn example on silhouette analysis. For each value it fits KMeans with random_state=10, prints the average silhouette_score, and draws one band per cluster from silhouette_samples. Most of the example's lines only draw the plots.

The video's make_blobs data

python
from sklearn.datasets import make_blobs

# 500 points, 2 features, 4 centres inside the box -10 to 10
X, y = make_blobs(n_samples=500, n_features=2, centers=4, cluster_std=1,
                  center_box=(-10.0, 10.0), shuffle=True, random_state=1)
range_n_clusters = [2, 3, 4, 5, 6]   # the cluster counts to compare

Scoring one value of n_clusters

python
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_samples, silhouette_score

clusterer = KMeans(n_clusters=4, random_state=10)
cluster_labels = clusterer.fit_predict(X)
silhouette_avg = silhouette_score(X, cluster_labels)              # mean s(i)
sample_silhouette_values = silhouette_samples(X, cluster_labels)  # one s(i) per point

Silhouette scores for 2 to 6 clusters

Reading the silhouette scores and plots · from the Complete Machine Learning in 6 Hours video · 331:30 to 334:55

For 2, 3 and 4 clusters the run below prints the video's numbers to every digit. For 5 and 6 clusters the video printed 0.56376 and 0.45047; this run prints 0.56146 and 0.48576. KMeans has changed since the video was recorded: since scikit-learn 1.4 it runs the algorithm once by default (n_init="auto" with the k-means++ start) instead of ten times, and other parts of the implementation have changed too, so for 5 and 6 clusters it settles on a slightly different grouping.

ExampleFrom the video, run on scikit-learn 1.9.1
for n_clusters in range_n_clusters:
    clusterer = KMeans(n_clusters=n_clusters, random_state=10)
    cluster_labels = clusterer.fit_predict(X)
    silhouette_avg = silhouette_score(X, cluster_labels)
    print("For n_clusters =", n_clusters, "The average silhouette_score is :", silhouette_avg)
    sample_silhouette_values = silhouette_samples(X, cluster_labels)
    print("    points below 0:", (sample_silhouette_values < 0).sum(),
          "  lowest s(i):", round(sample_silhouette_values.min(), 3))

Plotting the silhouette bands for 3 and 4 clusters

The scikit-learn example draws one figure per value of n_clusters. This shorter version draws the two cases the video compares most closely side by side: each coloured band is one cluster's sorted s(i) values, and the red dashed line is the average.

ExampleFrom the video, run on scikit-learn 1.9.1
import matplotlib.pyplot as plt
import matplotlib.cm as cm

fig, axes = plt.subplots(1, 2, figsize=(12, 5))
for ax, n_clusters in zip(axes, [3, 4]):
    cluster_labels = KMeans(n_clusters=n_clusters, random_state=10).fit_predict(X)
    silhouette_avg = silhouette_score(X, cluster_labels)
    sample_silhouette_values = silhouette_samples(X, cluster_labels)
    y_lower = 10
    for i in range(n_clusters):
        # one band per cluster: its s(i) values, sorted
        ith = np.sort(sample_silhouette_values[cluster_labels == i])
        y_upper = y_lower + len(ith)
        color = cm.nipy_spectral(float(i) / n_clusters)
        ax.fill_betweenx(np.arange(y_lower, y_upper), 0, ith, facecolor=color, edgecolor=color, alpha=0.7)
        ax.text(-0.05, y_lower + 0.5 * len(ith), str(i))
        y_lower = y_upper + 10
    ax.axvline(x=silhouette_avg, color="red", linestyle="--")
    ax.set_xlim([-0.1, 1])
    ax.set_yticks([])
    ax.set_title(f"n_clusters = {n_clusters}, average {silhouette_avg:.3f}")
    ax.set_xlabel("The silhouette coefficient values")
    ax.set_ylabel("Cluster label")
    print(n_clusters, "clusters, band widths:", np.bincount(cluster_labels))
plt.suptitle("Silhouette analysis for KMeans clustering on sample data")
plt.show()
Silhouette bands for three and four clusters: with three clusters the large cluster has a thin tail below zero; with four clusters the four bands are similar in width and stay at or above zero.

Why the video picks K = 4

  • K = 2 scores highest, 0.705, with no point below 0: the distinct blob is one cluster and the three close blobs share the other.
  • K = 3 has 5 points below 0, the tail the video circles at −0.1: they sit nearer another cluster than their own, so K = 3 is rejected.
  • K = 4 scores 0.651, and only one point dips below 0, to −0.009, too thin to see on the plot. Its four bands hold 123 to 128 points each.
  • K = 5 and 6 score lower and have points below 0 again.
  • Between the two clean options, 2 and 4, the video takes the larger one, and K = 4 agrees with the elbow method.

Validating hierarchical clustering on the Iris data

The silhouette score works the same way for Hierarchical clustering. The hierarchical clustering notebook in the course materials ends with this loop: ward clusters on the four scaled Iris features for 2 to 10 clusters, one silhouette_score each, then a line plot. Like the K-means silhouette loop, it starts at 2, because the score needs at least two clusters.

ExampleFrom the course notebook, run on scikit-learn 1.9.1
import pandas as pd
import matplotlib.pyplot as plt
from sklearn import datasets
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import AgglomerativeClustering
from sklearn.metrics import silhouette_score

iris = datasets.load_iris()
X_scaled = StandardScaler().fit_transform(pd.DataFrame(iris.data, columns=iris.feature_names))

silhouette_coefficients = []
for k in range(2, 11):
    agglo = AgglomerativeClustering(n_clusters=k, metric="euclidean", linkage="ward")
    agglo.fit(X_scaled)
    silhouette_coefficients.append(silhouette_score(X_scaled, agglo.labels_))
print([round(s, 3) for s in silhouette_coefficients])

plt.plot(range(2, 11), silhouette_coefficients)
plt.xticks(range(2, 11))
plt.xlabel("Number of Clusters")
plt.ylabel("Silhouette Coefficient")
plt.show()
A line plot of the silhouette coefficient against the number of ward clusters on the Iris data: highest at 2 clusters, lower at 3, and flat and low from 5 clusters on.

What the Iris scores show

  • 2 clusters score highest, 0.577, the choice the notebook and the dendrogram made.
  • 3 clusters score 0.447, although Iris has three species: versicolor and virginica overlap, so a third cluster splits them along a fuzzy border.
  • From 6 clusters on the score stays near 0.31: extra clusters cut real groups into pieces.

Silhouette score vs the elbow method

Elbow methodSilhouette score
MeasuresWCSS: how tight each cluster istightness and separation from the nearest cluster
Readinglook for the bend in the curvehigher is better, from −1 to +1
Per pointnoyes: silhouette_samples gives every s(i)
Used forpicking Kvalidating K and the clusters

Where you use the silhouette score

  • Choosing K for K-means, next to the elbow curve.
  • Comparing two clustering models on the same data, such as K-means against hierarchical clustering.
  • Finding misplaced points: a negative s(i) marks a point that sits nearer another cluster.
Watch out. The silhouette score favours round, well separated clusters. On the 250 scaled moon points in DBSCAN, the true two moons score 0.385, while K-means, which cuts across both moons, scores 0.495. Look at the plot as well as the number.
Try it yourself
  • In the nine-point example, change the label of [6, 2] from 1 to 0 and read its silhouette_samples value.
  • Change range_n_clusters to [2, 3, 4, 5, 6, 7, 8] in the score loop and see whether any K beats 0.705.
  • In the score loop, add n_init=10 to KMeans and compare the 5- and 6-cluster scores with the run above: ten starts settle on yet another grouping, and neither run reproduces the video's 0.564 and 0.450.

You understood something today that you didn't yesterday.