Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Logistic regression in scikit-learn

Logistic regression in scikit-learn is the LogisticRegression classifier, which you train with fit, tune with GridSearchCV and judge on a test set with a confusion matrix and a classification report.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

The theory lessons built the sigmoid, Log loss and the metrics of Precision, recall and F-beta. Here they run together on a real data set: is a breast tumour malignant or benign?

Loading the breast cancer data · from the Complete Machine Learning in 6 Hours video · 165:24 to 168:08

Loading the breast cancer data

load_breast_cancer ships with scikit-learn: 569 tumours, 30 independent features (mean radius, mean texture, mean perimeter, mean area and so on) and a target. In this data set 0 means malignant and 1 means benign. As in the video, the features go into a DataFrame with the feature names as columns, and value_counts checks whether the target is balanced.

python
import pandas as pd
from sklearn.datasets import load_breast_cancer

df = load_breast_cancer()
X = pd.DataFrame(df["data"], columns=df["feature_names"])    # 30 independent features
y = pd.Series(df["target"], name="Target")                    # 0 = malignant, 1 = benign
print(y.value_counts())

The video prints 357 ones and 212 zeros and treats it as balanced. It builds y as a one-column DataFrame; scikit-learn then prints a DataConversionWarning on fit, which a Series avoids.

Penalty, C and the parameter grid · from the Complete Machine Learning in 6 Hours video · 168:08 to 171:33

Choosing C, the penalty and the grid

The split is the one from Train and test split: test_size=0.33 and random_state=42, which leaves 188 tumours for testing. The LogisticRegression docs page shows the parameters that matter most:

  • The penalty, L1 or L2. The same two regularizations as in Lasso regression and Ridge regression, added to the log loss.
  • C, the inverse of regularization strength. Roughly 1 / λ: a smaller C means a stronger penalty.
  • class_weight. For an imbalanced data set, class_weight="balanced" gives the rare class more weight.

The docs page in the video is scikit-learn 1.0.2, where L1 is penalty='l1'. In 1.9.1 penalty is deprecated (removal is planned for 1.10) and the penalty is set with l1_ratio: 0 for L2, 1 for L1. L1 needs solver='liblinear' or 'saga', because the default lbfgs solver supports only L2.

The video's search: params = [{'C': [1, 5, 10]}, {'max_iter': [100, 150]}], a starting model LogisticRegression(C=100, max_iter=100), and GridSearchCV with scoring='f1' and cv=5. It scores with F1 because it is not sure whether false positives or false negatives matter more. Note the list of two dictionaries: GridSearchCV searches C alone, then max_iter alone, not the six combinations; a single dictionary would search all six.

Best parameters and the test report · from the Complete Machine Learning in 6 Hours video · 171:33 to 173:57

Running the video's search unscaled

"A lot of warnings will be coming": the run below repeats the video's search on unscaled features and counts the warnings instead of printing them all.

ExampleFrom the video, run on scikit-learn 1.9.1
import pandas as pd
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import confusion_matrix, accuracy_score

df = load_breast_cancer()
X = pd.DataFrame(df["data"], columns=df["feature_names"])
y = pd.Series(df["target"], name="Target")
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.33, random_state=42)

params = [{"C": [1, 5, 10]}, {"max_iter": [100, 150]}]
model1 = LogisticRegression(C=100, max_iter=100)
model = GridSearchCV(model1, param_grid=params, scoring="f1", cv=5)

import warnings
with warnings.catch_warnings(record=True) as caught:
    warnings.simplefilter("always")
    model.fit(X_train, y_train)
print(len(caught), "warnings, the first:", str(caught[0].message).splitlines()[0])

print(model.best_params_, round(model.best_score_, 4))
y_pred = model.predict(X_test)
print(confusion_matrix(y_test, y_pred))
print("accuracy:", round(accuracy_score(y_test, y_pred), 4))

What the unscaled run printed

  • All 26 warnings are ConvergenceWarning. lbfgs stops at max_iter before it converges, because the 30 features sit on very different scales (mean area in the hundreds, smoothness around 0.1).
  • The same winner as the video. best_params_ is {'max_iter': 150}; best_score_ is 0.9599 here against 0.9575 in the video.
  • A near copy of the video's test result. The confusion matrix is [[64, 3], [3, 118]] and accuracy 0.9681; the video printed [[63, 4], [3, 118]] and 0.9628. The video ran scikit-learn 1.0.2, and small differences like these are expected across versions.

Scaling inside a Pipeline

Pipeline with StandardScaler

Putting StandardScaler and LogisticRegression in one pipeline fixes the warnings, and GridSearchCV fits the scaler on each training fold only, so the validation fold stays unseen. Grid keys take the form <step name>__<parameter>; make_pipeline names each step after its class in lower case.

python
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

pipe = make_pipeline(StandardScaler(), LogisticRegression())
# one dictionary: all 3 x 2 = 6 combinations are searched
params = {"logisticregression__C": [1, 5, 10], "logisticregression__max_iter": [100, 150]}

Tuning and testing the scaled model

This run continues from the data and the split above.

ExampleRun on scikit-learn 1.9.1
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import classification_report

pipe = make_pipeline(StandardScaler(), LogisticRegression())
params = {"logisticregression__C": [1, 5, 10], "logisticregression__max_iter": [100, 150]}
model = GridSearchCV(pipe, param_grid=params, scoring="f1", cv=5)
model.fit(X_train, y_train)
print(model.best_params_, round(model.best_score_, 4))

y_pred = model.predict(X_test)
print(confusion_matrix(y_test, y_pred))
print("accuracy:", round(accuracy_score(y_test, y_pred), 4))
print(classification_report(y_test, y_pred))

Reading the scaled model's report

  • No warnings, and C = 1 wins. With scaled features lbfgs converges well inside 100 iterations, so max_iter 100 and 150 tie and the first value is kept. The cross-validated F1 rises to 0.9811.
  • The test matrix is [[66, 1], [3, 118]]. The top row is the 67 malignant tumours: 66 found, 1 called benign. The bottom row is the 121 benign ones: 118 right, 3 called malignant.
  • Accuracy 0.9787, up from 0.9681 unscaled, on the same 188 test tumours.
  • The report gives precision, recall and F1 per class. Class 0 (malignant) has recall 0.99, the share of malignant tumours found, the number a cancer screen cares about most. Since the data is fairly balanced, all the scores are high, as the video says.

L1 penalty with l1_ratio

Continuing from the same split, asking for L1 with the default solver fails:

ExampleThe default lbfgs solver with an L1 penalty
LogisticRegression(l1_ratio=1).fit(X_train, y_train)

With solver="liblinear" the L1 penalty works and, as with Lasso, sets some coefficients to exactly 0.

ExampleRun on scikit-learn 1.9.1
import numpy as np

for C in [0.1, 1]:
    l1_model = make_pipeline(StandardScaler(), LogisticRegression(l1_ratio=1, solver="liblinear", C=C, random_state=0))
    l1_model.fit(X_train, y_train)
    coef = l1_model[-1].coef_[0]
    print(f"C={C}: {np.sum(coef == 0)} of 30 coefficients are 0, test accuracy {l1_model.score(X_test, y_test):.4f}")

At C = 0.1 the L1 model keeps only 8 of the 30 features and still reaches 0.9681 test accuracy; at C = 1 it keeps 14 and reaches 0.9734, close to the L2 pipeline's 0.9787.

Unscaled vs scaled logistic regression

Unscaled (the video's code)StandardScaler in a Pipeline
Warnings on fitConvergenceWarningnone
best_params_max_iter 150C 1, max_iter 100
Cross-validated F10.95990.9811
Test confusion matrix[[64, 3], [3, 118]][[66, 1], [3, 118]]
Test accuracy0.96810.9787

Where you use LogisticRegression

  • A first model for any binary problem. It trains in milliseconds and sets the score other models must beat.
  • When probabilities matter. predict_proba gives a risk score you can threshold lower for a screening test.
  • When you want few features. The L1 penalty keeps a short list, as with Lasso for regression.
Watch out. In load_breast_cancer, 1 means benign. scoring="f1", precision and recall all treat 1 as the positive class by default, so they measure finding benign tumours. To score how well malignant tumours are found, pass pos_label=0, for example recall_score(y_test, y_pred, pos_label=0).
Try it yourself
  • Add class_weight="balanced" to the LogisticRegression in the pipeline and compare the confusion matrix.
  • Add 0.01 and 0.1 to the C values in params and see which C wins.
  • Print recall_score(y_test, y_pred, pos_label=0) (import it from sklearn.metrics) for the scaled model: the share of malignant tumours it finds.

Little by little, you're building something great.