Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Supervised and unsupervised learning

Supervised learning is a type of machine learning that trains on data with a known output column, while unsupervised learning finds structure in data that has no output column.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Most business problems fall into one of these two. Knowing which one you have tells you which family of algorithms to reach for, and how to check the result.

Learning from an output column

Supervised learning and regression · from the Complete Machine Learning in 6 Hours video · 7:54 to 12:29

The majority of business use cases fall into supervised or unsupervised machine learning. Supervised learning solves two kinds of problem, regression and classification. Unsupervised learning solves two others, clustering and dimensionality reduction. A third type, reinforcement learning, is named in the video and left for later; this course does not cover it.

Machine learning splits into supervised learning (regression and classification), unsupervised learning (clustering and dimensionality reduction) and reinforcement learning, which the video only names.

Take the video's dataset of age and weight: 24 and 62, 25 and 63, 21 and 72, 27 and 62. The task is to train a model on this data so that, given a new age, it outputs a weight. That trained model is also called the hypothesis.

  • Independent features are the inputs the model trains on. Here: age.
  • The dependent feature is the output to predict. Here: weight. It is called dependent because it changes when the input changes.
  • A supervised problem has one dependent feature and any number of independent features.
The board's age and weight table (24 and 62, 25 and 63, 21 and 72, 27 and 62) feeds a hypothesis that takes a new age and outputs a weight; age is the independent feature and weight the dependent feature.

Regression: a continuous output

The video's second table has ages 24, 23 and 25 with weights 72, 71 and 71.5. The output is a continuous number, so this is a regression problem. Plot the data as a scatter, draw a straight line y = mx + c through it, and read a new age's predicted weight off the line. That line is linear regression, the first algorithm of the course.

The notes for this lesson use a house price table instead: a house of size 5000 with 5 rooms sells for 450K, one of size 6000 with 6 rooms for 500K. Size and rooms are the independent features, price the dependent one, and price is continuous, so this is regression too.

A regression table of age and weight next to a scatter of age against weight with a red straight line y = mx + c; a new age is read off the line as a predicted weight.

Classifying, clustering and reducing dimensions

Classification, clustering and dimensionality reduction · from the Complete Machine Learning in 6 Hours video · 12:35 to 16:59

Classification: fixed categories

The independent features are the number of study hours, play hours and sleeping hours. The dependent feature is pass or fail. When the output has a fixed number of categories, the problem is classification: two categories make it binary classification, more than two make it multiclass classification. The board leaves the feature cells blank. The notes fill in two rows, with study and play hours: 7 hours of study and 3 of play is a pass, 2 of study and 6 of play is a fail. They also add a third answer, "may be", which turns the binary problem into a multiclass one.

A table with study hours and play hours as independent features and pass or fail as the dependent feature: 7 and 3 is a pass, 2 and 6 is a fail; two classes make it binary classification, and a third class, may be, makes it multiclass.

Clustering: grouping without an output

Now a dataset of salary and age with no output variable and no dependent variable. Clustering finds groups of similar people. In the video's example three clusters appear: people who are young with a high salary, people who are older with a good salary, and a middle-class group whose salary does not rise much with age.

This is customer segmentation. A company launching product 1 for rich customers and product 2 for middle-class customers can target each product's ads at the matching cluster. Later, regression or classification can run on each segment. The key word is grouping: clustering is not classification, because there is no output feature to learn.

The notes give a second version: an e-commerce company with each customer's salary and a spending score from 1 to 10, such as 20000 and 9, or 45000 and 2. Plotted, the customers fall into clusters, and the company can email each cluster its own discount.

Salary and age data with no output column, grouped into three clusters: young with a high salary, older with a good salary, and a middle-class group; product 1 is advertised to the rich clusters and product 2 to the middle-class one.

Dimensionality reduction: fewer features

With 1000 features, can the data be squeezed into fewer dimensions, say 100 features, while keeping most of the information? Dimensionality reduction algorithms such as PCA do that. The video also names LDA here. PCA ignores labels, so it is unsupervised. LDA (linear discriminant analysis) needs the class labels to find its directions, so it is a supervised method, even though it also reduces dimensions.

Dimensionality reduction: an algorithm such as PCA turns 1000 features into 100 features.

Sorting the algorithms by type

The notes close the topic with the algorithms sorted by type. The supervised ones from decision tree onwards solve both classification and regression:

TypeAlgorithms in the notesAlso in this course
Supervised, regressionLinear regression, ridge and lasso, ElasticNetPolynomial regression
Supervised, classificationLogistic regressionNaive Bayes
Supervised, bothDecision tree, random forest, AdaBoost, XGBoostK nearest neighbours, SVM
UnsupervisedK-means, hierarchical clustering, DBSCAN clusteringPCA, silhouette score

Reinforcement learning, the third type, has its own heading in the notes and no algorithms under it; this course does not cover it.

Fitting with and without an output in scikit-learn

scikit-learn shows the difference in one argument. A supervised model is trained with fit(X, y): features and the output. An unsupervised model is trained with fit(X): features only. To run this, install the libraries first (Installing scikit-learn).

The supervised fit

python
import numpy as np
from sklearn.linear_model import LinearRegression

age = np.array([[24], [25], [21], [27]])     # independent feature, as a column
weight = np.array([62, 63, 72, 62])          # dependent feature
model = LinearRegression().fit(age, weight)  # fit(X, y)

The unsupervised fit

python
from sklearn.cluster import KMeans

# age and salary (in thousands), no output column
people = np.array([[23, 90], [24, 95], [45, 120], [50, 130], [35, 30], [38, 35]])
kmeans = KMeans(n_clusters=3, n_init=10, random_state=0)
groups = kmeans.fit_predict(people)          # fit(X) only, then a group per row

Running both fits

ExampleAge and weight from the video; the six salary rows are made up for the clustering half
print("predicted weight for a new age of 23:", round(model.predict([[23]])[0], 1))
print("cluster of each person:", groups)
print("number of clusters found:", len(set(groups)))

Reading the two results

  • The supervised model returns a number. It learned from the weights in the table, so it can predict one for an age it has not seen. Simple linear regression explains the line behind it.
  • The clustering returns group numbers. The two young high earners share one number, the two older high earners another, and the two low earners a third. The numbers are names, not ranks: 0 is not better than 2.
  • Nothing in the clustering was labelled. The groups come only from how close the rows are to each other. K-means clustering covers how.

Supervised vs unsupervised learning

SupervisedUnsupervised
Output columnYes, the dependent featureNone
ProblemsRegression (continuous), classification (categories)Clustering, dimensionality reduction
Example from the videoAge to weight; study hours to pass or failSalary and age segments; 1000 to 100 features
scikit-learn callfit(X, y)fit(X)
Algorithms in this courseLinear and logistic regression, Naive Bayes, KNN, trees, ensembles, SVMK-means, hierarchical clustering, DBSCAN, PCA

Where you use supervised and unsupervised learning

  • Predicting a number such as a house price or a delivery time: supervised regression.
  • Predicting a category such as spam or not spam, pass or fail: supervised classification.
  • Customer segmentation for targeted ads, as in the video: unsupervised clustering, often followed by a supervised model per segment.
Watch out. Clustering is not classification. A classifier learns categories that already exist in the output column; clustering has no output column and makes up its own groups. The video's algorithm list also puts K nearest neighbours under clustering: KNN is a supervised classifier and regressor that needs labelled outputs.
Try it yourself
  • Change the new age from 23 to 30. The line slopes down in this tiny table, so the prediction drops.
  • Set n_clusters=2 and see which two groups merge.
  • Add a seventh person, [25, 92], to people and check that it joins the young high earners.

You understood something today that you didn't yesterday.