Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

One-vs-rest logistic regression

One-vs-rest (OvR) is a multiclass classification method that trains one binary logistic regression per class, each separating that class from all the others, and predicts the class whose model gives the highest probability.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Logistic regression draws one decision boundary between two classes. Many problems have three or more: a flower species, a product category, a digit. One-vs-rest, also called one-vs-all, turns such a problem into several binary ones that logistic regression already solves.

Splitting three classes into three binary problems

One-vs-rest models M1, M2 and M3 · from the Logistic Regression Multiclass Classification video · 0:28 to 3:26

With two classes, one best fit line separates class A from class B. Now take three classes. One-vs-rest treats one class as the positive category and groups the other two into one negative category, and trains a binary model M1 on that split. The next round makes the second class positive and the other two negative, model M2, and the third round gives M3. Three classes, three binary models.

Say the data has the input features f1, f2 and f3 and an output with three classes O1, O2 and O3. For the three models the output column becomes three columns, one per class: a row of class O1 is positive in the O1 column and negative in the other two. The video writes them as +1 and −1; the table here uses 1 and 0.

Three classes of points with three dividing lines, and a table whose output column O1, O2, O3, O1, O3, O2 becomes three 0 or 1 columns; model M1 learns the O1 column, M2 the O2 column and M3 the O3 column, and for a new point they give 0.25, 0.20 and 0.55, so the prediction is O3.

Predicting with the highest probability

Predicting with the highest probability · from the Logistic Regression Multiclass Classification video · 3:26 to 6:28

The video passes multi_class='ovr' to LogisticRegression; LogisticRegression in scikit-learn 1.9.1 has no multi_class parameter, and the code here uses OneVsRestClassifier.

M1 is trained on all the input features with the O1 column as its output, M2 with the O2 column and M3 with the O3 column. For new test data the features go to all three models, and each gives the probability that the point is its class: M1 0.25, M2 0.20 and M3 0.55. The highest is 0.55, from M3, so the output is O3, category 3. In the video M1 gives 0.20 and M2 0.25; either way M3's 0.55 is the highest.

Here the three probabilities add up to 1. Three separate models do not have to agree like that; scikit-learn divides each one by their sum before it reports them, and the largest still wins.

The three output columns and the highest probability

ExampleRun on scikit-learn 1.9.1
import pandas as pd

output = pd.Series(["O1", "O2", "O3", "O1", "O3", "O2"])   # the output column of the table
columns = pd.get_dummies(output).astype(int)                 # one 0/1 column per class
print(columns)

probs = {"M1": 0.25, "M2": 0.20, "M3": 0.55}                # each model's P(its class)
best = max(probs, key=probs.get)
print("highest:", best, probs[best], "-> output O" + best[1])

Training one-vs-rest in scikit-learn

The data: make_classification builds 1,000 rows with 10 features, 3 of them informative, and 3 classes, about a third each. 70% trains the models and 30% tests them.

The liblinear solver on three classes

The liblinear solver only fits binary models, and on three classes it says so:

Exampleliblinear on three classes
LogisticRegression(solver="liblinear").fit(X_train, y_train)

Older scikit-learn set one-vs-rest with LogisticRegression(multi_class='ovr'); LogisticRegression in 1.9.1 has no multi_class parameter, and the error names the replacement, OneVsRestClassifier.

OneVsRestClassifier

python
from sklearn.multiclass import OneVsRestClassifier
from sklearn.linear_model import LogisticRegression

ovr = OneVsRestClassifier(LogisticRegression())   # one binary model per class
ovr.fit(X_train, y_train)
print(len(ovr.estimators_))                        # 3 models: M1, M2, M3
print(ovr.predict_proba(X_test[:1]))               # one probability per class

Three classes from make_classification

The run fits the wrapper, scores it, rebuilds the three models by hand and compares with a single LogisticRegression.

ExampleRun on scikit-learn 1.9.1
import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression

X, y = make_classification(n_samples=1000, n_features=10, n_informative=3, n_classes=3, random_state=15)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.30, random_state=42)

from sklearn.multiclass import OneVsRestClassifier
from sklearn.metrics import accuracy_score, confusion_matrix, classification_report

ovr = OneVsRestClassifier(LogisticRegression()).fit(X_train, y_train)
y_pred = ovr.predict(X_test)
print("models:", len(ovr.estimators_), " accuracy:", round(accuracy_score(y_test, y_pred), 4))
print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred))

# the same three models by hand: class k against the rest, then the highest probability
raw = np.column_stack([LogisticRegression().fit(X_train, y_train == k).predict_proba(X_test)[:, 1] for k in range(3)])
print("row 0, raw P from M1, M2, M3:", raw[0].round(3), " sum", raw[0].sum().round(3))
print("row 0, ovr.predict_proba:    ", ovr.predict_proba(X_test[:1])[0].round(3))
print("by hand, same class on every row:", bool(np.all(raw.argmax(axis=1) == y_pred)))

multi = LogisticRegression().fit(X_train, y_train)          # one multinomial model
print("multinomial accuracy:", round(accuracy_score(y_test, multi.predict(X_test)), 4))

What the three models predicted

  • Three models, accuracy 0.79. ovr.estimators_ holds M1, M2 and M3, and 237 of the 300 test rows get the right class.
  • Class 1 is the hardest. Its row of the matrix is [3, 74, 25]: 25 rows of class 1 are predicted as class 2, so its recall is 0.73, against 0.82 for the other two.
  • The raw probabilities do not add up to 1. For the first test row M1, M2 and M3 give 0.000, 0.214 and 0.983, a sum of 1.196. predict_proba divides by that sum and reports 0.000, 0.179 and 0.821; the highest is still the third class.
  • The hand-built models agree. Taking the highest of the three probabilities gives the same class as OneVsRestClassifier on every test row.
  • A single multinomial model scores 0.7833. On this data the two approaches are almost level.

One-vs-rest vs multinomial logistic regression

One-vs-restMultinomial
Models trainedone binary model per classone model for all classes
Probabilitiesone sigmoid per class, divided by their sumone softmax over all classes
In scikit-learn 1.9.1OneVsRestClassifier(LogisticRegression())LogisticRegression() on 3 or more classes
Works with liblinearyesno
Accuracy on this data0.790.7833

Where you use one-vs-rest

  • Any binary classifier on many classes. OneVsRestClassifier wraps logistic regression, a linear SVM or any model that only separates two classes.
  • Reading one class at a time. Each of the models in estimators_ has its own coefficients, which show what separates that class from the rest.
  • Multi-label data. When a row can belong to several classes at once, each binary model answers its own yes or no.
Watch out. Pass the true labels first: confusion_matrix(y_test, y_pred). With the arguments swapped the 3 × 3 grid is transposed and the per-class precision and recall in classification_report trade places, which is easy to miss with three classes.
Try it yourself
  • Wrap the solver that failed: OneVsRestClassifier(LogisticRegression(solver="liblinear")) fits without an error.
  • Change n_classes=3 to n_classes=4 and range(3) to range(4): len(ovr.estimators_) becomes 4.
  • Print confusion_matrix(y_pred, y_test) and find where the 25 class-1 rows predicted as class 2 moved.

You understood something today that you didn't yesterday.