Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Logistic regression

Logistic regression is a classification algorithm that passes a straight line θ₀ + θ₁x through the sigmoid function, so its output is a probability between 0 and 1 that a 0.5 threshold turns into a class.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

The regression lessons predict a number. Many questions have a fixed set of answers instead: pass or fail, spam or not, cancer or not. Logistic regression is the first classification algorithm in the video, and it works very well for binary classification, two classes.

Pass or fail from study hours · from the Complete Machine Learning in 6 Hours video · 93:10 to 97:29

Predicting pass or fail with a straight line

The example: from a child's number of study hours (and play hours) predict whether the child passes or fails. There are two fixed categories, so this is binary classification. It can be extended to more classes, multiclass classification, but the video stays with two.

Keep one feature, the number of study hours. The board sketches the points without numbers; the table used here gives them hours: the students who studied 2, 3 and 4 hours failed, those who studied 5, 6 and 7 hours passed. Fail is written as 0 and pass as 1, and these are the training points.

Can linear regression solve it? It can draw a best fit line through the points, then use one rule: if hθ(x) is below 0.5 the output is 0, fail; if it is 0.5 or more the output is 1, pass. A new point predicted at 0.25 is a fail. A point far to the right lands above 0.5 on the extended line, a pass. So far it works.

An outlier breaks the straight line · from the Complete Machine Learning in 6 Hours video · 97:29 to 100:23

Breaking the line with one outlier

Now add an outlier: a student who studied 12 hours and, of course, passed. The best fit line moves to reach it: it tilts and flattens. The point where it crosses 0.5 moves to the right, so a student who studied a little more than the old threshold is now predicted to fail, although by the old line the student passes. One outlier changed the answer for other students.

The second problem: the line gives values above 1 on the right and, projected back to the left, negative values. The answer can only be between 0 and 1. The fix for both is to squash the line, with the sigmoid function.

Study hours on the x axis and fail or pass on the y axis: a green straight line fitted to six students crosses 0.5 at 4.50 hours; adding one passing student at 12 hours flattens the grey line so it crosses 0.5 at 4.96 hours, and both lines go below 0 and above 1.

Fitted on these points, the line crosses 0.5 at exactly 4.5 hours, halfway between the last fail and the first pass. With the 12-hour student it crosses at 4.96 hours, so a student who studied 4.6 hours flips from pass to fail.

The decision boundary and the sigmoid · from the Complete Machine Learning in 6 Hours video · 100:23 to 105:10

Squashing the line with the sigmoid

The hypothesis starts as a line: hθ(x) = θ₀ + θ₁x₁ + θ₂x₂ + … + θₙxₙ, written short as θᵀx. With one feature it is θ₀ + θ₁x. On top of this value logistic regression applies a function g that squashes it: hθ(x) = g(θ₀ + θ₁x). Let z = θ₀ + θ₁x. Then hθ(x) = g(z), with

This g is the sigmoid or logistic function. Its graph is an S: close to 0 on the far left, 1 on the far right, and exactly 0.5 at z = 0. Many texts write it σ(z) and call it the sigmoid activation; g and σ are the same function.

The S shaped sigmoid curve g(z) rises from 0 to 1 and passes 0.5 at z equal to 0, so the prediction is 1 whenever z is 0 or more; a card gives z as theta zero plus theta one x and g of z as 1 over 1 plus e to the minus z.

The key reading of the curve: g(z) ≥ 0.5 whenever z ≥ 0, and g(z) < 0.5 whenever z < 0. So the model predicts pass exactly when z = θ₀ + θ₁x is 0 or more. The line θ₀ + θ₁x = 0 is the decision boundary. The name comes from the two halves: regression builds the straight line, the logistic function squashes it. Does the squashing help with the outlier problem? The video says yes; the code checks it.

Fitting the study hours data

The LogisticRegression class

predict_proba returns the sigmoid output for each class, predict applies the 0.5 threshold. intercept_ is θ₀ and coef_ is θ₁, so the boundary z = 0 sits at x = −θ₀ / θ₁.

python
from sklearn.linear_model import LogisticRegression

model = LogisticRegression()
model.fit(hours, result)                 # hours: one column, result: 0 or 1
print(model.predict_proba([[4.6]]))      # [P(fail), P(pass)]
print(model.predict([[4.6]]))            # 1 when P(pass) >= 0.5

Linear and logistic with and without the outlier

ExampleFrom the video's board, run on scikit-learn 1.9.1
import numpy as np
from sklearn.linear_model import LinearRegression, LogisticRegression

hours = np.array([2, 3, 4, 5, 6, 7]).reshape(-1, 1)
result = np.array([0, 0, 0, 1, 1, 1])                 # 0 = fail, 1 = pass
hours_out = np.vstack([hours, [[12]]])                       # the outlier: 12 hours, pass
result_out = np.append(result, 1)

for name, X, y in [("6 students", hours, result), ("with the 12 h outlier", hours_out, result_out)]:
    lin = LinearRegression().fit(X, y)
    cut = (0.5 - lin.intercept_) / lin.coef_[0]              # where the line crosses 0.5
    at_46 = lin.predict([[4.6]])[0]
    log = LogisticRegression().fit(X, y)
    boundary = -log.intercept_[0] / log.coef_[0, 0]          # where z = 0
    p_46 = log.predict_proba([[4.6]])[0, 1]
    print(f"{name}:")
    print(f"  linear:   0.5 at {cut:.2f} h, 4.6 h -> {at_46:.2f}")
    print(f"  logistic: boundary at {boundary:.2f} h, 4.6 h -> P(pass) {p_46:.2f}")

lin = LinearRegression().fit(hours, result)
print("linear output at 0 h and 10 h:", lin.predict([[0], [10]]).round(2))

Plotting the sigmoid

ExampleRun on scikit-learn 1.9.1
import numpy as np
import matplotlib.pyplot as plt

def g(z):
    return 1 / (1 + np.exp(-z))

for z in [-4, -1, 0, 1, 4]:
    print(f"g({z:2}) = {g(z):.3f}")

z = np.linspace(-6, 6, 200)
plt.plot(z, g(z), color="red")
plt.axhline(0.5, color="gray", linestyle="--")
plt.axvline(0, color="gray", linestyle="--")
plt.title("Sigmoid function")
plt.xlabel("z")
plt.ylabel("g(z)")
plt.show()
A red S shaped sigmoid curve from 0 to 1, crossing the dashed 0.5 line exactly where z is 0.

What the fits and the curve show

  • The outlier moves the linear threshold. The 0.5 point slides from 4.50 to 4.96 hours, and the student at 4.6 hours drops from 0.53 (pass) to 0.46 (fail).
  • The outlier does not move logistic regression. Its boundary stays at 4.50 hours and P(pass) at 4.6 hours stays 0.53. A point far on the correct side sits on the flat part of the S and barely pulls on the fit.
  • The line leaves the 0 to 1 range. At 0 hours it predicts −0.66 and at 10 hours 1.91, neither of which is a valid answer.
  • The sigmoid stays inside it. g(−4) = 0.018, g(0) = 0.5 and g(4) = 0.982: every output is a probability, and the 0.5 line meets the curve at z = 0.

Logistic regression vs linear regression

Linear regressionLogistic regression
Predictsa numbera class, through a probability
Output rangeany value, below 0 and above 1between 0 and 1
Hypothesisθ₀ + θ₁xg(θ₀ + θ₁x), g = sigmoid
Decisionclass 1 when g(z) ≥ 0.5, which is z ≥ 0
Effect of a far outliertilts the line and moves the thresholdalmost none

Where you use logistic regression

  • Yes or no questions. Pass or fail, spam or not, a loan repaid or not, a tumour malignant or benign.
  • When you need a probability. predict_proba gives a number you can rank or threshold, not only a label.
  • As the first classifier to try. It is fast, and its coefficients show which way each feature pushes the answer.
Watch out. Despite the name, logistic regression is a classifier. predict returns classes 0 and 1, never the probability; ask predict_proba for that. And 0.5 is a choice, not a law: a cancer screen may flag a patient at a lower probability.
Try it yourself
  • Move the outlier from 12 to 20 hours. The linear 0.5 point moves on to 5.28 hours, so now even a student at 5 hours, who passed, is predicted to fail; the logistic boundary stays at 4.50.
  • Print g(10) and g(-10): the sigmoid gets very close to 1 and 0 but never reaches them.
  • Fit LogisticRegression on the 6 students and print predict_proba([[4.5]]). On the boundary the two probabilities are both about 0.5.

You understood something today that you didn't yesterday.