Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Support vector regression (SVR)

Support vector regression (SVR) is the regression form of the support vector machine that fits a line with a tube of width ε around it and counts only the points outside the tube as errors.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Support vector machines (SVM) separated two classes with the widest margin. SVR turns the margin around: the points should fall inside it. Linear regression pays for every residual, however small; SVR pays nothing for an error smaller than ε.

Drawing the ε tube around the line

The worked example predicts a house price from its size. SVR draws the best fit line wᵀx + b, the same hyperplane equation as the classifier, and two marginal lines parallel to it, wᵀx + b + ε above and wᵀx + b − ε below. ε (epsilon) is the margin of error: a prediction within ε of the true price costs nothing.

Some points fall outside the tube. For each of them, ξᵢ (xi, the slack) is its distance past the edge of the tube, the error above the margin. The points on or outside the tube are the support vectors: they decide where the line goes, and the points inside the tube do not.

Price against size with a regression line w transpose x plus b and two dashed lines at plus and minus epsilon around it, forming a tube; green points inside the tube cost nothing, and the circled red points outside it are support vectors, each with a slack distance xi to the edge of the tube.

Writing the SVR cost function

The cost function has the same two parts as the soft margin SVM: keep w small, and pay C for every unit of slack. The constraint says each point is within ε of the line, plus its own slack ξᵢ.

The board writes ‖w‖/2 and |yᵢ − wᵢxᵢ|; the standard form squares the norm, as in the classifier, and keeps b inside the absolute value. Each point's slack is its ε-insensitive loss, zero inside the tube and growing in a straight line outside it. It plays the part the hinge loss plays for the classifier.

C is a hyperparameter. A larger C makes slack more expensive, so the line bends toward the points outside the tube and the loss on the training data falls; too large a C fits noise.

Fitting SVR on ten house sizes

Ten houses, size in hundreds of square feet and price in lakhs, with a tube of ε = 3.

A linear SVR with ε = 3

kernel="linear" gives a straight line. epsilon is the half width of the tube, in the units of the target. After fitting, coef_ holds w, intercept_ holds b and support_ the row numbers of the support vectors.

python
from sklearn.svm import SVR

svr = SVR(kernel="linear", C=10, epsilon=3).fit(size, price)
residual = price - svr.predict(size)
loss = np.maximum(0, np.abs(residual) - 3)          # the epsilon-insensitive loss
ExampleRun on scikit-learn 1.9.1
import numpy as np
from sklearn.svm import SVR

size = np.array([1, 2, 3, 4, 5, 6, 7, 8, 9, 10]).reshape(-1, 1)    # hundreds of square feet
price = np.array([12, 15, 23, 19, 26, 30, 28, 36, 34, 45])         # lakhs

svr = SVR(kernel="linear", C=10, epsilon=3).fit(size, price)
print("w:", svr.coef_[0].round(3), "  b:", svr.intercept_.round(3))
residual = price - svr.predict(size)
print("residuals:", residual.round(2))
print("support vectors (house numbers):", svr.support_ + 1)
print("epsilon-insensitive loss:", np.maximum(0, np.abs(residual) - 3).round(2))

What the tube fit shows

  • The line is price = 3 × size + 10 (w = 3, b = 10).
  • Five houses are support vectors, numbers 3, 4, 7, 9 and 10: their residuals, 4, −3, −3, −3 and 5, are at least ε = 3 in size. Houses 4, 7 and 9 sit exactly on the edge of the tube.
  • Only houses 3 and 10 have a loss, 1 and 2, the part of their error beyond ε. The five houses inside the tube cost nothing and do not shape the line.

Predicting the total bill on the tips dataset

The tips dataset has 244 restaurant bills with the tip, the customer's sex, smoker or not, the day, lunch or dinner, and the size of the party. The task is to predict the total bill. sns.load_dataset downloads it once.

Encoding the columns and fitting SVR

The four text columns become 0 or 1 columns with pd.get_dummies. SVR() uses the rbf kernel by default, with C=1 and epsilon=0.1.

python
df = sns.load_dataset("tips")
X = pd.get_dummies(df[["tip", "sex", "smoker", "day", "time", "size"]], drop_first=True,
                   columns=["sex", "smoker", "day", "time"]).astype(float)
y = df["total_bill"]
svr = SVR().fit(X_train, y_train)

Tuning C and gamma with GridSearchCV

gamma sets how far one point's influence reaches in the rbf kernel. Hyperparameter tuning with GridSearchCV tries every pair from the grid with 5-fold cross-validation.

python
param_grid = {"C": [0.1, 1, 10, 100, 1000], "gamma": [1, 0.1, 0.01, 0.001, 0.0001], "kernel": ["rbf"]}
grid = GridSearchCV(SVR(), param_grid, cv=5).fit(X_train, y_train)
grid.best_params_
ExampleThe tips dataset, run on scikit-learn 1.9.1
import pandas as pd
import seaborn as sns
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.svm import SVR
from sklearn.metrics import r2_score, mean_absolute_error

df = sns.load_dataset("tips")
X = pd.get_dummies(df[["tip", "sex", "smoker", "day", "time", "size"]], drop_first=True,
                   columns=["sex", "smoker", "day", "time"]).astype(float)
y = df["total_bill"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=10)

svr = SVR().fit(X_train, y_train)
pred = svr.predict(X_test)
print(f"default SVR:  R2 {r2_score(y_test, pred):.3f}   MAE {mean_absolute_error(y_test, pred):.3f}")
print("support vectors:", len(svr.support_), "of", len(X_train))

param_grid = {"C": [0.1, 1, 10, 100, 1000], "gamma": [1, 0.1, 0.01, 0.001, 0.0001], "kernel": ["rbf"]}
grid = GridSearchCV(SVR(), param_grid, cv=5).fit(X_train, y_train)
pred = grid.predict(X_test)
print("best:", grid.best_params_)
print(f"tuned SVR:    R2 {r2_score(y_test, pred):.3f}   MAE {mean_absolute_error(y_test, pred):.3f}")

What the tips models show

  • The default SVR scores a test R² of 0.462, with a mean absolute error of 4.117 dollars.
  • 180 of the 183 training bills are support vectors: the default ε = 0.1 dollars is so narrow that almost every bill falls outside the tube.
  • The grid picks C = 1000 and gamma = 0.0001, a smooth rbf kernel with expensive errors, and the test R² rises to 0.508 with an error of 3.868 dollars.

SVR vs linear regression

SVRLinear regression
Which errors costonly those larger than εevery residual
How an error costsin a straight line past ε (ε-insensitive loss)squared
Points that shape the fitthe support vectors, on or outside the tubeevery training point
Curved relationshipskernels (rbf, polynomial)only by adding features by hand
Settings to tuneC, epsilon, kernel and gammanone for plain least squares

Where you use support vector regression

  • Small and medium tabular datasets, such as the 244 tips, where a curved fit helps and training time is not a problem.
  • Data with a few large errors, which the straight-line loss outside the tube punishes less than squared error does.
  • A known tolerance: when an error of up to ε does not matter, such as a price quoted to the nearest lakh.
Watch out. epsilon is measured in the target's units, and its default of 0.1 is tiny for a bill in dollars and meaningless for a price in rupees. Pick ε from the size of error you can accept, and tune C and gamma: on the tips data the defaults gave a lower R² than the grid's choice.
Try it yourself
  • Change epsilon=3 to epsilon=6 on the ten houses and count the support vectors.
  • Set C=0.01 on the ten houses and compare w with the run above.
  • Fit SVR(kernel="linear") on the tips data and compare its test R² with the tuned rbf model.

Slow is fine. Stopping is the only problem.