Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

SVM kernels

An SVM kernel is a function that lets a support vector machine separate data that no straight line can split, by working as if the points had been moved into a higher dimension.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Support vector machines (SVM) drew a straight hyperplane. When one class sits inside the other, every straight line fails, and a kernel is the way out.

Data that is not linearly separable · from the SVM kernels in-depth intuition video · 4:22 to 7:21

Moving the points into a higher dimension

The video starts from a dataset that is not linearly separable: one class in the middle, the other around it. Any best fit line leaves many points on the wrong side. An SVM kernel converts the data from two dimensions into three: the middle points are raised, the outer points stay low, and a plane divides them.

The transformation from a lower dimension to a higher one is done with a mathematical formula, and each kind of kernel has its own: polynomial, rbf (radial basis function) and sigmoid.

Squaring one feature to separate it · from the SVM kernels in-depth intuition video · 7:21 to 10:52

Squaring one feature to separate one dimension

The simplest case has one feature x. The points of one class sit near 0, the points of the other class on both sides of them. A single cut on x leaves errors whichever way it goes. The video applies y = f(x) = x² to every point (the data is centred at 0). The points now lie on a parabola: the middle class low, the outer class high, and a straight line separates them.

The same idea works in any number of dimensions: two dimensions go to three, three to four, and so on.

Left: red points near 0 and green points on both sides of them lie on one line, so no single cut separates them. Right: after the transformation y equals x squared the points sit on a parabola, the red ones low and the green ones high, and a horizontal straight line splits them.

Adding x² as a second feature

ExampleRun on scikit-learn 1.9.1
import numpy as np
from sklearn.svm import SVC

x = np.array([-3.4, -3.0, -2.6, -2.2, -1.2, -0.9, -0.5, -0.2, 0.2, 0.6, 0.9, 1.3, 2.2, 2.6, 3.0, 3.5])
y = np.array([0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0])     # 1 = the middle class

one_d = x.reshape(-1, 1)
two_d = np.c_[x, x ** 2]                                            # add y = x squared
print("linear SVC on x:        accuracy", SVC(kernel="linear").fit(one_d, y).score(one_d, y))
print("linear SVC on x and x^2: accuracy", SVC(kernel="linear").fit(two_d, y).score(two_d, y))

The linear SVC on x alone gets 0.6875 of the 16 points right (11 of them), the best a single cut can do. With x² added as a second feature it gets every point right.

The polynomial kernel and its features · from the SVM kernels in-depth intuition video · 11:26 to 15:14

Expanding the polynomial kernel

For two features x₁ and x₂ the video writes the polynomial kernel as the vector product of the points plus 1, raised to the power d. Working out the product gives the terms x₁², x₂² and x₁x₂ next to x₁ and x₂. Keeping each term once, the two dimensions become five features: x₁, x₂, x₁², x₂² and x₁x₂, and in that space a plane can split the classes.

A kernel takes two points, not one. For d = 2 and c = 1, (aᵀb + 1)² equals the dot product of the two points after each is mapped to (x₁², x₂², √2·x₁x₂, √2·x₁, √2·x₂, 1). The SVM only ever needs those dot products, so the kernel gives the answer without building the new columns. That shortcut is called the kernel trick.

Splitting the new features with a plane · from the SVM kernels in-depth intuition video · 15:14 to 18:06

Training a linear SVC on the polynomial features

Plotted with x₁², x₂² and x₁x₂ as the axes, the points of the two classes separate, so a linear SVC finds a plane between them. Selecting kernel="poly" does the same work internally, and the choice between the polynomial, rbf and sigmoid kernels is a hyperparameter, tuned like any other.

The example takes 100 heights between −5 and 5. At each height it places a point on a circle of radius 5 on both sides, which draws the whole inner circle, and a point on a circle of radius 10 on both sides, which draws two arcs to the left and right of it. The inner circle is class 1, the arcs class 0.

The circle and the two arcs

python
x = np.linspace(-5.0, 5.0, 100)                    # 100 heights between -5 and 5
outer = np.sqrt(10 ** 2 - x ** 2)                  # on a circle of radius 10: two arcs
inner = np.sqrt(5 ** 2 - x ** 2)                   # on a circle of radius 5: the whole circle
X = np.vstack([np.c_[np.hstack([outer, -outer]), np.hstack([x, -x])],
               np.c_[np.hstack([inner, -inner]), np.hstack([x, -x])]])
y = np.r_[np.zeros(200), np.ones(200)]             # 0 = the arcs, 1 = the inner circle

The polynomial features by hand

python
Xp = np.c_[X, X[:, 0] ** 2, X[:, 1] ** 2, X[:, 0] * X[:, 1]]    # x1, x2, x1^2, x2^2, x1*x2
SVC(kernel="linear").fit(Xp_train, y_train).score(Xp_test, y_test)
ExampleA circle and two arcs, run on scikit-learn 1.9.1
import numpy as np
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.svm import SVC

x = np.linspace(-5.0, 5.0, 100)                    # 100 heights between -5 and 5
outer = np.sqrt(10 ** 2 - x ** 2)                  # on a circle of radius 10: two arcs
inner = np.sqrt(5 ** 2 - x ** 2)                   # on a circle of radius 5: the whole circle
X = np.vstack([np.c_[np.hstack([outer, -outer]), np.hstack([x, -x])],
               np.c_[np.hstack([inner, -inner]), np.hstack([x, -x])]])
y = np.r_[np.zeros(200), np.ones(200)]             # 0 = the arcs, 1 = the inner circle
Xp = np.c_[X, X[:, 0] ** 2, X[:, 1] ** 2, X[:, 0] * X[:, 1]]

X_train, X_test, Xp_train, Xp_test, y_train, y_test = train_test_split(X, Xp, y, test_size=0.25, random_state=0)
print("linear SVC on x1, x2:              ", SVC(kernel="linear").fit(X_train, y_train).score(X_test, y_test))
print("linear SVC on the 5 poly features: ", SVC(kernel="linear").fit(Xp_train, y_train).score(Xp_test, y_test))
for kernel, extra in [("poly", {}), ("poly", {"degree": 2, "coef0": 1}), ("rbf", {}), ("sigmoid", {})]:
    acc = SVC(kernel=kernel, **extra).fit(X_train, y_train).score(X_test, y_test)
    print(f"kernel={kernel:8} {str(extra):27} on x1, x2: {acc:.3f}")

fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 4))
ax1.scatter(X[:, 0], X[:, 1], c=y, cmap="coolwarm", s=12)
ax1.set_title("A circle between two arcs: no straight line splits them")
ax1.set_xlabel("x1"); ax1.set_ylabel("x2")
ax2.scatter(X[:, 0] ** 2, X[:, 1] ** 2, c=y, cmap="coolwarm", s=12)
ax2.set_title("The same points on x1² and x2²")
ax2.set_xlabel("x1²"); ax2.set_ylabel("x2²")
plt.show()
Left: the inner circle of radius 5 in red between two blue arcs of radius 10, plotted in x1 and x2, which no straight line separates. Right: the same points against x1 squared and x2 squared, where the circle becomes the straight line x1 squared plus x2 squared equals 25 and the arcs the line equals 100, so a straight line splits them.

What the circles show

  • A linear SVC on x₁ and x₂ scores 0.45, worse than guessing, because the arcs lie on both sides of the circle.
  • On the five polynomial features the same linear SVC scores 1.0. In the right-hand plot the circle lies on the line x₁² + x₂² = 25 and the arcs on x₁² + x₂² = 100, so a straight line splits them.
  • kernel="poly" with its defaults scores 0.590; with degree=2 and coef0=1, the video's formula, it scores 1.000. The rbf kernel scores 1.000 too, and the sigmoid kernel 0.510.

Choosing a kernel and its settings

On a less tidy dataset the kernels differ less, and the settings matter as much as the kernel. 1,000 points from make_classification, two features and two clusters per class, are fitted with each of the four kernels, then the rbf kernel is tuned over C and gamma.

ExampleRun on scikit-learn 1.9.1
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.svm import SVC

X, y = make_classification(n_samples=1000, n_features=2, n_classes=2, n_clusters_per_class=2,
                           n_redundant=0, random_state=1)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=10)
for kernel in ("linear", "rbf", "poly", "sigmoid"):
    print(f"{kernel:8} test accuracy {SVC(kernel=kernel).fit(X_train, y_train).score(X_test, y_test):.3f}")

param_grid = {"C": [0.1, 1, 10, 100, 1000], "gamma": [1, 0.1, 0.01, 0.001, 0.0001], "kernel": ["rbf"]}
grid = GridSearchCV(SVC(), param_grid, cv=5).fit(X_train, y_train)
print("best:", grid.best_params_, f" test accuracy {grid.score(X_test, y_test):.3f}")

What the four kernels scored

  • The rbf kernel scores highest, 0.852, then linear 0.832, polynomial 0.812 and sigmoid 0.772.
  • The grid picks C = 1 and gamma = 1 and scores 0.852 again: on this data the tuned settings do no better than the defaults.

Polynomial kernel vs rbf kernel

Polynomial kernelrbf kernel
Formula(γ·aᵀb + coef0)^degreee^(−γ‖a − b‖²)
New features it stands forevery product of the features up to the degreeinfinitely many; a similarity that fades with distance
Settingsdegree, coef0, gamma, Cgamma, C
Good forboundaries of a known shape, such as circles (degree 2)a default when the shape is unknown
scikit-learnSVC(kernel="poly")SVC(kernel="rbf"), the default

Where you use SVM kernels

  • Curved class boundaries on small and medium datasets, where a linear SVM underfits.
  • Image and text classification with a few thousand rows, such as handwritten digits, where an rbf SVM is a strong baseline.
  • Support vector regression: Support vector regression (SVR) takes the same kernels.
Watch out. scikit-learn's kernel="poly" defaults to degree=3 and coef0=0, which drops the lower terms of the video's (aᵀb + 1)d. On the circles that kernel scored far below the degree=2, coef0=1 version. Set coef0 when you want the + 1.
Try it yourself
  • Remove the x ** 2 column from two_d and use np.c_[x, np.abs(x)] instead: does a linear SVC still separate the classes?
  • Change {"degree": 2, "coef0": 1} to {"degree": 2, "coef0": 0} on the circles and compare the accuracy.
  • Change random_state=1 to random_state=0 in make_classification and check whether the four kernels keep their order.

This is what real progress feels like.