Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Instance-based vs model-based learning

Instance-based learning is a way of learning in which the model memorises the training data and answers a new case from the stored examples closest to it, while model-based learning generalises the data into a pattern, such as a formula or a decision boundary, and answers from that pattern.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Supervised and unsupervised learning describe the data. This split describes what a model keeps after training: the data itself, or a pattern learned from it. It decides how much a saved model weighs and how fast it answers.

Memorising versus generalising

The notes start from a use case and an ML model that solves a regression or a classification problem. How the model learns the pattern of the data can go two ways, the same two ways a person learns:

  • Memorising gives instance-based learning: keep the data, and answer from it directly. The flow is data, then output.
  • Generalising gives model-based learning: find the pattern in the data, turn it into a generalized model, and answer from the model. The flow is data, then pattern, then generalization method.
Two panels of study hours against play hours with pass and fail points. Instance-based learning keeps the data and answers a new student from the stored points closest to it, as KNN does. Model-based learning learns a pattern from the data, a decision boundary with pass on one side and fail on the other, and predicts from that boundary.

Predicting from the nearest students

The notes' example is a table with the number of play hours, the number of study hours and pass or not. Plot study hours against play hours, each student a cross. A new student arrives: the test point.

Instance-based learning draws a circle around the test point and looks at the training points inside it. If most of those students passed, the answer is pass. Nothing was learned in advance; the comparison happens when the question arrives, and every question needs the stored data. K nearest neighbours (KNN) works this way, and K nearest neighbours (KNN) covers it in full. The notes compare it to a domain expert who answers from the cases they remember.

Drawing a decision boundary

Model-based learning studies the same points once, during training, and learns their pattern: here a curve, the decision boundary, with pass on one side and fail on the other. A new student is placed on one side of the curve and gets that side's answer. Once the boundary is learned, the training rows are no longer needed; the model is the handful of numbers that draw the curve. The straight line of Simple linear regression, with its two numbers θ₀ and θ₁, is a model-based pattern too.

Saving a trained model to disk

The notes then follow what happens after training. The model is saved to the hard drive in a serialized format, such as a pickle (.pkl) file of a few kilobytes or megabytes. Later the saved file is loaded, takes a new input and returns its output. For a model-based model the file holds the learned numbers. An instance-based model has no pattern to save, so its file has to carry the training data, and it grows with every row.

The notes summarise the difference in a table, from preparing the data to scoring a new case:

StepModel-based (usual) learningInstance-based learning
Preparing the dataPrepare the data for trainingThe same: no difference here
TrainingTrain a model to estimate its parameters: the pattern is discoveredNo training: discovering the pattern waits until a query arrives
After trainingStore the model; the training data can be thrown awayThere is no model to store; the training data must be kept
GeneralisingRules are generalised into a model before any new case is seenNo generalisation in advance; each new case is handled as it comes
PredictingUse the modelUse the training data directly, part or all of it
StorageUsually smallUsually large
Scoring a new caseUsually fastCan be slow

Measuring what each model stores

KNN is instance-based and logistic regression is model-based (a classifier that learns a straight boundary, taught in Logistic regression). Fit both on the same students, then pickle them the way a saved model is written to disk and count the bytes.

The study and play hours data

The video gives no rows for this example, so the students here are generated: a student who studies more than they play usually passes.

python
import numpy as np

rng = np.random.default_rng(0)

def students(n):
    # study and play hours for n students; more study than play tends to pass
    study, play = rng.uniform(0, 10, n), rng.uniform(0, 10, n)
    passed = (study - play + rng.normal(0, 1, n) > 0).astype(int)   # 1 = pass, 0 = fail
    return np.column_stack([study, play]), passed

Fitting KNN and logistic regression

python
from sklearn.neighbors import KNeighborsClassifier
from sklearn.linear_model import LogisticRegression

X, y = students(100)
knn = KNeighborsClassifier(n_neighbors=5).fit(X, y)   # instance-based: stores X and y
logreg = LogisticRegression().fit(X, y)                # model-based: learns a boundary

Comparing the saved sizes as the data grows

ExampleGenerated students, run on scikit-learn 1.9.1
import pickle

new_student = [[7, 3]]   # 7 study hours, 3 play hours
print("KNN prediction:    ", knn.predict(new_student)[0])
print("logistic prediction:", logreg.predict(new_student)[0])

for n in (100, 10_000):
    X, y = students(n)
    knn_bytes = len(pickle.dumps(KNeighborsClassifier(n_neighbors=5).fit(X, y)))
    log_bytes = len(pickle.dumps(LogisticRegression().fit(X, y)))
    print(f"{n:>6} students: KNN file {knn_bytes:>7} bytes, logistic file {log_bytes} bytes")

print("numbers inside the logistic model:", logreg.coef_.size + logreg.intercept_.size)

Reading the sizes

  • Both predict 1, a pass, for 7 study hours and 3 play hours, the notes' passing row. They agree here; they get there differently.
  • The KNN file grows with the data. It holds every training row, so 100 times more students makes a file about 80 times bigger, and each prediction compares the new student with the stored rows.
  • The logistic file stays the same size. Its pattern is three numbers, two coefficients and an intercept, whether it learned from 100 students or 10,000.

Instance-based vs model-based learning

Instance-basedModel-based
Learns byMemorising the training dataGeneralising it into a pattern
Work done at training timeAlmost none: store the rowsMost of it: estimate the parameters
Work done at prediction timeCompare with the stored rowsApply the formula
Saved modelGrows with the dataA fixed set of numbers
Examples in this courseK nearest neighboursLinear and logistic regression, trees, SVM

Where you use instance-based and model-based learning

  • Instance-based for small tables where similar past cases are the best guide, and new rows keep arriving: adding a row to KNN needs no retraining.
  • Model-based for most production systems: a small saved file and fast predictions, such as a pickled regression model behind an app.
  • Interviews: "what is the difference between instance-based and model-based learning?" is answered by memorising versus generalising, with KNN as the example.
Watch out. An instance-based model is slow to predict on a big table, because every prediction searches the stored rows. A model that takes seconds per answer on 10 million rows is usually a KNN-style model, not a slow computer. If you need fast answers, choose a model-based algorithm or shrink the stored data.
Try it yourself
  • Change (100, 10_000) to (100, 100_000): the KNN file grows again and the logistic file does not.
  • Predict for [[2, 6]], the notes' failing student, with both models.
  • Set n_neighbors=1: KNN now copies the single closest student, the purest form of memorising.

Slow is fine. Stopping is the only problem.