Installing Python for statistics
A Python statistics stack is a set of libraries, NumPy, SciPy, pandas, statsmodels, matplotlib and seaborn, that stores data, computes summaries and statistical tests, and draws charts.
Last updated: 07 Oct, 2026 · SciPy 1.18
Every lesson runs on the same six libraries. NumPy holds arrays of numbers, pandas holds tables, SciPy computes distributions and tests, statsmodels adds more tests and models, matplotlib draws charts and seaborn draws statistical charts and loads example datasets. There are no API keys anywhere. You can install them on your own computer (Windows, macOS or Linux) or use Google Colab in a browser.
Checking your Python version
The stack needs Python 3.12 or newer: NumPy 2.5 and SciPy 1.18 do not install on older versions. Open a terminal (on Windows, PowerShell from the Start menu; on macOS, the Terminal app) and ask Python for its version:
python3 --versionIf it prints 3.12 or higher, you are set. If it prints 3.11 or lower, or the command is not found, either install a current Python from python.org (on Windows, tick Add python.exe to PATH in the installer), or take the uv route below, which downloads a suitable Python for the project by itself.
Installing the libraries
Pick one route. On your own computer, uv is the one to prefer: it makes a project folder with its own environment, so these pinned versions never clash with other Python work. pip into your main Python is quicker to type. Colab needs no local install. The six libraries are pinned so your output matches the lessons; JupyterLab is unpinned.
curl -LsSf https://astral.sh/uv/install.sh | sh
uv init stats-course --python 3.12
cd stats-course
uv add numpy==2.5.3 scipy==1.18.1 pandas==3.0.6 statsmodels==0.15.0 seaborn==0.13.2 matplotlib==3.11.2 jupyterlabWhat the uv route sets up
- The first line installs uv, a fast Python package manager. Open a new terminal afterwards so the
uvcommand is found. uv init stats-course --python 3.12creates the folderstats-coursewith apyproject.tomlfile that lists the project's libraries, and pins the project to Python 3.12. If your computer does not have that version, uv downloads it the first time it builds the environment.uv addcreates a private environment instats-course/.venv, installs the libraries into it and records the exact versions inuv.lock, so a teammate gets the same install withuv sync.- Run code with
uv run, for exampleuv run python check_setup.py. There is no activation step: uv always uses the project's own environment.
What the pip and Colab routes do
python3 -m pip(orpy -m pipon Windows) installs into the same Python that thepython3(orpy) command runs, which avoids installing into one Python and running another.- Colab comes with these libraries already, in its own versions. The
%pipline replaces them with the pinned ones. If pip prints a warning that another preinstalled Colab package expects a different version, the lessons are not affected, since they import only these six. Restart the session afterwards (Runtime, Restart session) so the new versions load.
Writing code in Jupyter or VS Code
The video's practicals run in a Jupyter notebook: code in cells, each cell's output and chart right under it. Start JupyterLab from the project folder and it opens in your browser:
uv run jupyter labIn JupyterLab, choose File, New, Notebook, pick the Python 3 kernel, and paste a lesson's code into a cell. To use VS Code instead, install its Python and Jupyter extensions, open the stats-course folder, and pick the interpreter with Python: Select Interpreter from the Command Palette: for the uv route, choose the one in .venv. In a notebook a chart appears under its cell; in a .py file it opens in a window when the code reaches plt.show(), which is why every plotting example ends with that line.
Checking the installed versions
Save this as check_setup.py in the project folder and run it with uv run python check_setup.py (uv), python3 check_setup.py or py check_setup.py (pip), or paste it into a notebook cell:
import sys
import numpy, scipy, pandas, statsmodels, matplotlib, seaborn
print("Python ", sys.version.split()[0])
print("numpy ", numpy.__version__)
print("scipy ", scipy.__version__)
print("pandas ", pandas.__version__)
print("statsmodels", statsmodels.__version__)
print("matplotlib ", matplotlib.__version__)
print("seaborn ", seaborn.__version__)Python 3.12.13 numpy 2.5.3 scipy 1.18.1 pandas 3.0.6 statsmodels 0.15.0 matplotlib 3.11.2 seaborn 0.13.2
Reading the version check
- Python 3.12 or newer is the floor for the whole stack.
- SciPy 1.18.1 is the version every test and distribution output in the lessons comes from.
- The other five match the pins in the install line. A different version can still run every lesson, but a few printed numbers or warnings may differ.
Loading data for the lessons
Most lessons type the video's numbers straight into the code. A few use a real dataset from seaborn.
A list of numbers from the board
ages = [10, 12, 14, 18, 24, 26, 30, 35, 36, 37, 40, 41, 42, 43, 50, 51] # the video's histogram agesA dataset from seaborn
sns.load_dataset downloads a small CSV file from the seaborn-data repository on GitHub the first time, then reads a cached copy (sns.get_data_home() prints the folder):
import seaborn as sns
tips = sns.load_dataset("tips") # restaurant bills: 244 rows
iris = sns.load_dataset("iris") # flower measurements: 150 rowsLoading the tips and iris datasets
print("ages:", len(ages), "values, from", min(ages), "to", max(ages))
print("tips:", tips.shape)
print(tips.head(3).to_string())
print("iris:", iris.shape)
print(iris["species"].value_counts().to_string())ages: 16 values, from 10 to 51 tips: (244, 7) total_bill tip sex smoker day time size 0 16.99 1.01 Female No Sun Dinner 2 1 10.34 1.66 Male No Sun Dinner 3 2 21.01 3.50 Male No Sun Dinner 3 iris: (150, 5) species setosa 50 versicolor 50 virginica 50
Reading the dataset shapes
- 16 values: the ages are a plain Python list; pandas and NumPy accept it as it is.
- (244, 7): tips has 244 restaurant bills and 7 columns, numbers such as
total_billand categories such asday. - (150, 5): iris has 150 flowers, four measurements each and a
speciescolumn with 50 flowers of each of three species.
Fixing a failed install
Most install problems print one of a few messages. Find yours:
python3: command not foundor'py' is not recognized: Python is not installed, or not on the PATH. Install it from python.org (tick Add python.exe to PATH on Windows), or use uv, which brings its own Python.No matching distribution found for numpy==2.5.3, after a list of older versions: the Python running pip is older than 3.12, so pip only sees older NumPy releases. Check the version as above, then use a 3.12 install oruv init --python 3.12.error: externally-managed-environment: this Python belongs to the operating system or Homebrew, which blocks pip from changing it. Use the uv route; it installs into the project's own environment.ModuleNotFoundError: No module named 'scipy': the code runs with a different Python from the one you installed into. In the uv project, run it withuv run; in VS Code, select the.venvinterpreter; in Colab, rerun the install cell after a restart.- A
URLErrorfromsns.load_dataset: the first load needs an internet connection. Connect and run it again; later runs read the cached copy. - A Colab cell still prints the old version: the session was not restarted after
%pip install. - A script runs but no chart appears: the plotting code is missing
plt.show()at the end.
uv vs pip vs Colab
| uv | pip | Colab | |
|---|---|---|---|
| Where it runs | Your computer | Your computer | A browser, on Google's machines |
| Install line | uv add ... inside a project | python3 -m pip install ... | %pip install ... in a cell |
| Keeps versions per project | Yes, in pyproject.toml and uv.lock | No, one shared Python | Per session; reinstall after a reset |
| Installs Python for you | Yes | No | Python is already there |
| Good for | Repeatable projects | A quick start | No local install |
Where you use each setup
- Following the lessons in a notebook or a
.pyfile: every example runs with these six libraries. - A data analysis of your own: the uv route records the exact versions, so the same numbers come out on a teammate's computer.
- A computer you cannot install on, such as a locked work laptop: Colab runs everything on Google's machines.
df.corr() raise an error on a table with text columns unless you pass numeric_only=True, and code from older tutorials stops there. When an output does not match a lesson, print the versions before anything else.Related
- Previous: Statistics
- Next: Descriptive and inferential statistics
- Reference: SciPy statistical functions and the uv documentation
- Print
tips.describe()to see the count, mean, smallest and largest value of each numeric column. - Print
sns.get_dataset_names()to list every dataset seaborn can load. - Run
import scipy.stats as st; print(st.norm.cdf(0)). It prints 0.5, a first taste of the normal distribution.
Little by little, you're building something great.