Installing scikit-learn
scikit-learn is a Python library that provides machine learning algorithms, datasets and evaluation tools behind one consistent fit and predict interface.
Last updated: 05 Oct, 2026 · scikit-learn 1.9.1
Every lesson runs on the same small stack: scikit-learn for the models, NumPy and pandas for the data, matplotlib for the plots, and XGBoost for the boosting lessons. There are no API keys anywhere in the course. You can install it on your own computer (Windows, macOS or Linux) or skip the install and use Google Colab in a browser.
Checking your Python version
The stack needs Python 3.12 or newer: NumPy 2.5, SciPy 1.18 and XGBoost 3.4 do not install on older versions. Open a terminal (on Windows, PowerShell from the Start menu; on macOS, the Terminal app) and ask Python for its version:
python3 --versionIf it prints 3.12 or higher, you are set. If it prints 3.11 or lower, or the command is not found, either install a current Python from python.org (on Windows, tick Add python.exe to PATH in the installer), or take the uv route below, which downloads a suitable Python for the project by itself.
Installing the libraries
Pick one route. On your own computer, the uv route is the one to prefer: it makes a project folder with its own environment, so these pinned versions never clash with other Python work. pip into your main Python is quicker to type. Colab needs no install at all. The course libraries are pinned so your output matches the lessons; JupyterLab is unpinned.
curl -LsSf https://astral.sh/uv/install.sh | sh
uv init ml-course --python 3.12
cd ml-course
uv add scikit-learn==1.9.1 xgboost==3.4.1 numpy==2.5.3 pandas==3.0.6 matplotlib==3.11.2 jupyterlabWhat the uv route sets up
- The first line installs uv, a fast Python package manager. Open a new terminal afterwards so the
uvcommand is found. uv init ml-course --python 3.12creates the folderml-coursewith apyproject.tomlfile that lists the project's libraries, and pins the project to Python 3.12, downloading it if your computer does not have it.uv addcreates a private environment inml-course/.venv, installs the libraries into it and records the exact versions inuv.lock, so a teammate gets the same install.- Run code with
uv run, for exampleuv run python check_setup.py. There is no activation step: uv always uses the project's own environment.
What the pip and Colab routes do
python3 -m pip(orpy -m pipon Windows) installs into the same Python that thepython3(orpy) command runs, which avoids the classic mistake of installing into one Python and running another.- Colab already has NumPy, pandas and matplotlib, so its line installs only the two pinned libraries. After a
%pip install, restart the session (Runtime, Restart session) so the new versions load, then run your cells again.
import xgboost fails with a message about libomp, install it with Homebrew: brew install libomp (Homebrew itself comes from brew.sh). Windows, Linux and Colab need nothing extra.Writing code in Jupyter or VS Code
The video writes its practicals in a Jupyter notebook: code in cells, each cell's output right under it. Start JupyterLab from the project folder and it opens in your browser:
uv run jupyter labIn JupyterLab, choose File, New, Notebook, pick the Python 3 kernel, and paste a lesson's code into a cell. To use VS Code instead, install its Python and Jupyter extensions, open the ml-course folder, and pick the interpreter with Python: Select Interpreter from the Command Palette: choose the one in .venv for the uv route. A .ipynb file then runs in VS Code with the same kernel, and a .py file runs with the Run button.
Checking the installed versions
Save this as check_setup.py in the project folder and run it with uv run python check_setup.py (uv), python3 check_setup.py or py check_setup.py (pip), or paste it into a notebook cell:
import sys
import sklearn, numpy, pandas, matplotlib, xgboost
print("Python ", sys.version.split()[0])
print("scikit-learn", sklearn.__version__)
print("numpy ", numpy.__version__)
print("pandas ", pandas.__version__)
print("matplotlib ", matplotlib.__version__)
print("xgboost ", xgboost.__version__)Python 3.12.13 scikit-learn 1.9.1 numpy 2.5.3 pandas 3.0.6 matplotlib 3.11.2 xgboost 3.4.1
Reading the version check
- Python 3.12 or newer is the floor for the whole stack.
- scikit-learn 1.9.1 is the version every output in these lessons comes from.
- xgboost printing a version means OpenMP loaded; on macOS a failure here is the libomp case above.
- The other three match the pins in the install line.
Loading the course datasets
The lessons use three kinds of data, and none needs a manual download:
A dataset inside the package
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True) # 150 flowers, 4 measurements eachA dataset fetched once
from sklearn.datasets import fetch_california_housing
housing = fetch_california_housing() # downloads on first call, then reads a local cacheA CSV file from the notes
The practical notebooks in the video's materials come with their CSV files. pandas reads one straight from its raw GitHub address:
import pandas as pd
url = ("https://raw.githubusercontent.com/krishnaik06/The-Grand-Complete-Data-Science-Materials/"
"main/Machine%20Learning/2-Complete%20Linear%20Regression/Practicals/height-weight.csv")
df = pd.read_csv(url) # read straight from GitHub, nothing to download firstLoading Iris, California housing and a CSV
print("iris:", X.shape, "labels:", sorted(set(y.tolist())))
print("california:", housing.data.shape)
print("features:", housing.feature_names)
print("height-weight csv:", df.shape)
print(df.head(3).to_string())iris: (150, 4) labels: [0, 1, 2] california: (20640, 8) features: ['MedInc', 'HouseAge', 'AveRooms', 'AveBedrms', 'Population', 'AveOccup', 'Latitude', 'Longitude'] height-weight csv: (23, 2) Weight Height 0 45 120 1 58 135 2 48 123
Reading the dataset shapes
- (150, 4): Iris has 150 rows and 4 feature columns, with three classes 0, 1 and 2.
- (20640, 8): California housing has 20,640 districts and 8 features. It replaces the Boston table the video uses, which scikit-learn removed in 1.2. The first call downloads it into a
scikit_learn_datafolder in your home directory, so later runs work offline. - (23, 2): the notes' height and weight table, read over the internet each time it runs. Save it with
df.to_csv("height-weight.csv", index=False)to work offline.
Fixing a failed install
Most install problems print one of a few messages. Find yours:
python3: command not foundor'py' is not recognized: Python is not installed, or not on the PATH. Install it from python.org (tick Add python.exe to PATH on Windows), or use uv, which brings its own Python.No matching distribution found for numpy==2.5.3, after a long list of older versions: the Python running pip is older than 3.12, so pip can only see older NumPy releases. Check the version as above, then use a 3.12 install oruv init --python 3.12.error: externally-managed-environment: this Python belongs to the operating system or Homebrew, which blocks pip from changing it. Use the uv route; it installs into the project's own environment.XGBoostError: XGBoost Library (libxgboost.dylib) could not be loadedon macOS, withLibrary not loaded: @rpath/libomp.dylibfurther down: runbrew install libomp.ModuleNotFoundError: No module named 'sklearn': the code runs with a different Python from the one you installed into. In the uv project, run it withuv run; in VS Code, select the.venvinterpreter; in Colab, rerun the install cell after a restart. The package installs asscikit-learnand imports assklearn;pip install sklearnis the wrong name.- A Colab cell still prints the old scikit-learn version: the session was not restarted after
%pip install. - A URL or download error from
fetch_california_housingorread_csv: the first load needs an internet connection; check it, then run again.
uv vs pip vs Colab
| uv | pip | Colab | |
|---|---|---|---|
| Where it runs | Your computer | Your computer | A browser, on Google's machines |
| Install line | uv add ... inside a project | python3 -m pip install ... | %pip install ... in a cell |
| Keeps versions per project | Yes, in pyproject.toml and uv.lock | No, one shared Python | Per session; reinstall after a reset |
| Installs Python for you | Yes | No | Python is already there |
| Good for | Repeatable projects | A quick start | No local install at all |
Where you use each setup
- Following the lessons in a notebook or a
.pyfile: every example runs with these five libraries. - A new project of your own: the uv route records the exact versions, so a teammate gets the same results with
uv sync. - A computer you cannot install on, such as a locked work laptop: Colab runs everything on Google's machines.
n_init default in 1.4. When an output does not match the lesson, print sklearn.__version__ before anything else.Related
- Previous: Instance-based vs model-based learning
- Next: Train and test split
- Reference: Installing scikit-learn and the uv documentation
- Run
import sklearn; sklearn.show_versions()to see every library scikit-learn depends on. - Print
housing.DESCR[:600]to read what each California housing column means. - Print
df.describe()for the height and weight table: the mean, smallest and largest value of each column.
This is what real progress feels like.