Machine Learningscikit-learn 1.9.1 · xgboost 3.4.1 · Python 3.12+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Installing scikit-learn

scikit-learn is a Python library that provides machine learning algorithms, datasets and evaluation tools behind one consistent fit and predict interface.

Last updated: 05 Oct, 2026 · scikit-learn 1.9.1

Every lesson runs on the same small stack: scikit-learn for the models, NumPy and pandas for the data, matplotlib for the plots, and XGBoost for the boosting lessons. There are no API keys anywhere in the course. You can install it on your own computer (Windows, macOS or Linux) or skip the install and use Google Colab in a browser.

Checking your Python version

The stack needs Python 3.12 or newer: NumPy 2.5, SciPy 1.18 and XGBoost 3.4 do not install on older versions. Open a terminal (on Windows, PowerShell from the Start menu; on macOS, the Terminal app) and ask Python for its version:

python3 --version

If it prints 3.12 or higher, you are set. If it prints 3.11 or lower, or the command is not found, either install a current Python from python.org (on Windows, tick Add python.exe to PATH in the installer), or take the uv route below, which downloads a suitable Python for the project by itself.

Installing the libraries

Pick one route. On your own computer, the uv route is the one to prefer: it makes a project folder with its own environment, so these pinned versions never clash with other Python work. pip into your main Python is quicker to type. Colab needs no install at all. The course libraries are pinned so your output matches the lessons; JupyterLab is unpinned.

curl -LsSf https://astral.sh/uv/install.sh | sh
uv init ml-course --python 3.12
cd ml-course
uv add scikit-learn==1.9.1 xgboost==3.4.1 numpy==2.5.3 pandas==3.0.6 matplotlib==3.11.2 jupyterlab

What the uv route sets up

  • The first line installs uv, a fast Python package manager. Open a new terminal afterwards so the uv command is found.
  • uv init ml-course --python 3.12 creates the folder ml-course with a pyproject.toml file that lists the project's libraries, and pins the project to Python 3.12, downloading it if your computer does not have it.
  • uv add creates a private environment in ml-course/.venv, installs the libraries into it and records the exact versions in uv.lock, so a teammate gets the same install.
  • Run code with uv run, for example uv run python check_setup.py. There is no activation step: uv always uses the project's own environment.

What the pip and Colab routes do

  • python3 -m pip (or py -m pip on Windows) installs into the same Python that the python3 (or py) command runs, which avoids the classic mistake of installing into one Python and running another.
  • Colab already has NumPy, pandas and matplotlib, so its line installs only the two pinned libraries. After a %pip install, restart the session (Runtime, Restart session) so the new versions load, then run your cells again.
macOS and XGBoost. XGBoost needs the OpenMP runtime, which macOS does not ship. If import xgboost fails with a message about libomp, install it with Homebrew: brew install libomp (Homebrew itself comes from brew.sh). Windows, Linux and Colab need nothing extra.

Writing code in Jupyter or VS Code

The video writes its practicals in a Jupyter notebook: code in cells, each cell's output right under it. Start JupyterLab from the project folder and it opens in your browser:

uv run jupyter lab

In JupyterLab, choose File, New, Notebook, pick the Python 3 kernel, and paste a lesson's code into a cell. To use VS Code instead, install its Python and Jupyter extensions, open the ml-course folder, and pick the interpreter with Python: Select Interpreter from the Command Palette: choose the one in .venv for the uv route. A .ipynb file then runs in VS Code with the same kernel, and a .py file runs with the Run button.

Checking the installed versions

Save this as check_setup.py in the project folder and run it with uv run python check_setup.py (uv), python3 check_setup.py or py check_setup.py (pip), or paste it into a notebook cell:

ExampleRun on the course's own install
import sys
import sklearn, numpy, pandas, matplotlib, xgboost

print("Python      ", sys.version.split()[0])
print("scikit-learn", sklearn.__version__)
print("numpy       ", numpy.__version__)
print("pandas      ", pandas.__version__)
print("matplotlib  ", matplotlib.__version__)
print("xgboost     ", xgboost.__version__)

Reading the version check

  • Python 3.12 or newer is the floor for the whole stack.
  • scikit-learn 1.9.1 is the version every output in these lessons comes from.
  • xgboost printing a version means OpenMP loaded; on macOS a failure here is the libomp case above.
  • The other three match the pins in the install line.

Loading the course datasets

The lessons use three kinds of data, and none needs a manual download:

A dataset inside the package

python
from sklearn.datasets import load_iris

X, y = load_iris(return_X_y=True)   # 150 flowers, 4 measurements each

A dataset fetched once

python
from sklearn.datasets import fetch_california_housing

housing = fetch_california_housing()   # downloads on first call, then reads a local cache

A CSV file from the notes

The practical notebooks in the video's materials come with their CSV files. pandas reads one straight from its raw GitHub address:

python
import pandas as pd

url = ("https://raw.githubusercontent.com/krishnaik06/The-Grand-Complete-Data-Science-Materials/"
       "main/Machine%20Learning/2-Complete%20Linear%20Regression/Practicals/height-weight.csv")
df = pd.read_csv(url)   # read straight from GitHub, nothing to download first

Loading Iris, California housing and a CSV

ExampleRun on scikit-learn 1.9.1
print("iris:", X.shape, "labels:", sorted(set(y.tolist())))
print("california:", housing.data.shape)
print("features:", housing.feature_names)
print("height-weight csv:", df.shape)
print(df.head(3).to_string())

Reading the dataset shapes

  • (150, 4): Iris has 150 rows and 4 feature columns, with three classes 0, 1 and 2.
  • (20640, 8): California housing has 20,640 districts and 8 features. It replaces the Boston table the video uses, which scikit-learn removed in 1.2. The first call downloads it into a scikit_learn_data folder in your home directory, so later runs work offline.
  • (23, 2): the notes' height and weight table, read over the internet each time it runs. Save it with df.to_csv("height-weight.csv", index=False) to work offline.

Fixing a failed install

Most install problems print one of a few messages. Find yours:

  • python3: command not found or 'py' is not recognized: Python is not installed, or not on the PATH. Install it from python.org (tick Add python.exe to PATH on Windows), or use uv, which brings its own Python.
  • No matching distribution found for numpy==2.5.3, after a long list of older versions: the Python running pip is older than 3.12, so pip can only see older NumPy releases. Check the version as above, then use a 3.12 install or uv init --python 3.12.
  • error: externally-managed-environment: this Python belongs to the operating system or Homebrew, which blocks pip from changing it. Use the uv route; it installs into the project's own environment.
  • XGBoostError: XGBoost Library (libxgboost.dylib) could not be loaded on macOS, with Library not loaded: @rpath/libomp.dylib further down: run brew install libomp.
  • ModuleNotFoundError: No module named 'sklearn': the code runs with a different Python from the one you installed into. In the uv project, run it with uv run; in VS Code, select the .venv interpreter; in Colab, rerun the install cell after a restart. The package installs as scikit-learn and imports as sklearn; pip install sklearn is the wrong name.
  • A Colab cell still prints the old scikit-learn version: the session was not restarted after %pip install.
  • A URL or download error from fetch_california_housing or read_csv: the first load needs an internet connection; check it, then run again.

uv vs pip vs Colab

uvpipColab
Where it runsYour computerYour computerA browser, on Google's machines
Install lineuv add ... inside a projectpython3 -m pip install ...%pip install ... in a cell
Keeps versions per projectYes, in pyproject.toml and uv.lockNo, one shared PythonPer session; reinstall after a reset
Installs Python for youYesNoPython is already there
Good forRepeatable projectsA quick startNo local install at all

Where you use each setup

  • Following the lessons in a notebook or a .py file: every example runs with these five libraries.
  • A new project of your own: the uv route records the exact versions, so a teammate gets the same results with uv sync.
  • A computer you cannot install on, such as a locked work laptop: Colab runs everything on Google's machines.
Watch out. A different scikit-learn version can print slightly different numbers. For example, KMeans changed its n_init default in 1.4. When an output does not match the lesson, print sklearn.__version__ before anything else.
Try it yourself
  • Run import sklearn; sklearn.show_versions() to see every library scikit-learn depends on.
  • Print housing.DESCR[:600] to read what each California housing column means.
  • Print df.describe() for the height and weight table: the mean, smallest and largest value of each column.

This is what real progress feels like.