StatisticsSciPy 1.18 · pandas 3.0 · statsmodels 0.15 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
57 small wins to finish your pathNext lesson →

Installing Python for statistics

A Python statistics stack is a set of libraries, NumPy, SciPy, pandas, statsmodels, matplotlib and seaborn, that stores data, computes summaries and statistical tests, and draws charts.

Last updated: 07 Oct, 2026 · SciPy 1.18

Every lesson runs on the same six libraries. NumPy holds arrays of numbers, pandas holds tables, SciPy computes distributions and tests, statsmodels adds more tests and models, matplotlib draws charts and seaborn draws statistical charts and loads example datasets. There are no API keys anywhere. You can install them on your own computer (Windows, macOS or Linux) or use Google Colab in a browser.

Checking your Python version

The stack needs Python 3.12 or newer: NumPy 2.5 and SciPy 1.18 do not install on older versions. Open a terminal (on Windows, PowerShell from the Start menu; on macOS, the Terminal app) and ask Python for its version:

python3 --version

If it prints 3.12 or higher, you are set. If it prints 3.11 or lower, or the command is not found, either install a current Python from python.org (on Windows, tick Add python.exe to PATH in the installer), or take the uv route below, which downloads a suitable Python for the project by itself.

Installing the libraries

Pick one route. On your own computer, uv is the one to prefer: it makes a project folder with its own environment, so these pinned versions never clash with other Python work. pip into your main Python is quicker to type. Colab needs no local install. The six libraries are pinned so your output matches the lessons; JupyterLab is unpinned.

curl -LsSf https://astral.sh/uv/install.sh | sh
uv init stats-course --python 3.12
cd stats-course
uv add numpy==2.5.3 scipy==1.18.1 pandas==3.0.6 statsmodels==0.15.0 seaborn==0.13.2 matplotlib==3.11.2 jupyterlab

What the uv route sets up

  • The first line installs uv, a fast Python package manager. Open a new terminal afterwards so the uv command is found.
  • uv init stats-course --python 3.12 creates the folder stats-course with a pyproject.toml file that lists the project's libraries, and pins the project to Python 3.12. If your computer does not have that version, uv downloads it the first time it builds the environment.
  • uv add creates a private environment in stats-course/.venv, installs the libraries into it and records the exact versions in uv.lock, so a teammate gets the same install with uv sync.
  • Run code with uv run, for example uv run python check_setup.py. There is no activation step: uv always uses the project's own environment.

What the pip and Colab routes do

  • python3 -m pip (or py -m pip on Windows) installs into the same Python that the python3 (or py) command runs, which avoids installing into one Python and running another.
  • Colab comes with these libraries already, in its own versions. The %pip line replaces them with the pinned ones. If pip prints a warning that another preinstalled Colab package expects a different version, the lessons are not affected, since they import only these six. Restart the session afterwards (Runtime, Restart session) so the new versions load.

Writing code in Jupyter or VS Code

The video's practicals run in a Jupyter notebook: code in cells, each cell's output and chart right under it. Start JupyterLab from the project folder and it opens in your browser:

uv run jupyter lab

In JupyterLab, choose File, New, Notebook, pick the Python 3 kernel, and paste a lesson's code into a cell. To use VS Code instead, install its Python and Jupyter extensions, open the stats-course folder, and pick the interpreter with Python: Select Interpreter from the Command Palette: for the uv route, choose the one in .venv. In a notebook a chart appears under its cell; in a .py file it opens in a window when the code reaches plt.show(), which is why every plotting example ends with that line.

Checking the installed versions

Save this as check_setup.py in the project folder and run it with uv run python check_setup.py (uv), python3 check_setup.py or py check_setup.py (pip), or paste it into a notebook cell:

ExampleRun on the course's own install
import sys
import numpy, scipy, pandas, statsmodels, matplotlib, seaborn

print("Python     ", sys.version.split()[0])
print("numpy      ", numpy.__version__)
print("scipy      ", scipy.__version__)
print("pandas     ", pandas.__version__)
print("statsmodels", statsmodels.__version__)
print("matplotlib ", matplotlib.__version__)
print("seaborn    ", seaborn.__version__)

Reading the version check

  • Python 3.12 or newer is the floor for the whole stack.
  • SciPy 1.18.1 is the version every test and distribution output in the lessons comes from.
  • The other five match the pins in the install line. A different version can still run every lesson, but a few printed numbers or warnings may differ.

Loading data for the lessons

Most lessons type the video's numbers straight into the code. A few use a real dataset from seaborn.

A list of numbers from the board

python
ages = [10, 12, 14, 18, 24, 26, 30, 35, 36, 37, 40, 41, 42, 43, 50, 51]   # the video's histogram ages

A dataset from seaborn

sns.load_dataset downloads a small CSV file from the seaborn-data repository on GitHub the first time, then reads a cached copy (sns.get_data_home() prints the folder):

python
import seaborn as sns

tips = sns.load_dataset("tips")   # restaurant bills: 244 rows
iris = sns.load_dataset("iris")   # flower measurements: 150 rows

Loading the tips and iris datasets

ExampleRun on seaborn 0.13.2
print("ages:", len(ages), "values, from", min(ages), "to", max(ages))
print("tips:", tips.shape)
print(tips.head(3).to_string())
print("iris:", iris.shape)
print(iris["species"].value_counts().to_string())

Reading the dataset shapes

  • 16 values: the ages are a plain Python list; pandas and NumPy accept it as it is.
  • (244, 7): tips has 244 restaurant bills and 7 columns, numbers such as total_bill and categories such as day.
  • (150, 5): iris has 150 flowers, four measurements each and a species column with 50 flowers of each of three species.

Fixing a failed install

Most install problems print one of a few messages. Find yours:

  • python3: command not found or 'py' is not recognized: Python is not installed, or not on the PATH. Install it from python.org (tick Add python.exe to PATH on Windows), or use uv, which brings its own Python.
  • No matching distribution found for numpy==2.5.3, after a list of older versions: the Python running pip is older than 3.12, so pip only sees older NumPy releases. Check the version as above, then use a 3.12 install or uv init --python 3.12.
  • error: externally-managed-environment: this Python belongs to the operating system or Homebrew, which blocks pip from changing it. Use the uv route; it installs into the project's own environment.
  • ModuleNotFoundError: No module named 'scipy': the code runs with a different Python from the one you installed into. In the uv project, run it with uv run; in VS Code, select the .venv interpreter; in Colab, rerun the install cell after a restart.
  • A URLError from sns.load_dataset: the first load needs an internet connection. Connect and run it again; later runs read the cached copy.
  • A Colab cell still prints the old version: the session was not restarted after %pip install.
  • A script runs but no chart appears: the plotting code is missing plt.show() at the end.

uv vs pip vs Colab

uvpipColab
Where it runsYour computerYour computerA browser, on Google's machines
Install lineuv add ... inside a projectpython3 -m pip install ...%pip install ... in a cell
Keeps versions per projectYes, in pyproject.toml and uv.lockNo, one shared PythonPer session; reinstall after a reset
Installs Python for youYesNoPython is already there
Good forRepeatable projectsA quick startNo local install

Where you use each setup

  • Following the lessons in a notebook or a .py file: every example runs with these six libraries.
  • A data analysis of your own: the uv route records the exact versions, so the same numbers come out on a teammate's computer.
  • A computer you cannot install on, such as a locked work laptop: Colab runs everything on Google's machines.
Watch out. Library versions change defaults. pandas 2.0 made df.corr() raise an error on a table with text columns unless you pass numeric_only=True, and code from older tutorials stops there. When an output does not match a lesson, print the versions before anything else.
Try it yourself
  • Print tips.describe() to see the count, mean, smallest and largest value of each numeric column.
  • Print sns.get_dataset_names() to list every dataset seaborn can load.
  • Run import scipy.stats as st; print(st.norm.cdf(0)). It prints 0.5, a first taste of the normal distribution.
PreviousStatistics

Little by little, you're building something great.