Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Installing Python for NLP

NLTK (Natural Language Toolkit) is a Python library of text-processing tools and language data: tokenizers, stemmers, a lemmatizer, stopword lists, a part-of-speech tagger and a named entity chunker.

Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9

The machine learning parts of the course run on a small stack: NLTK for preprocessing, scikit-learn for vectors and classifiers, gensim for Word2Vec, and NumPy, pandas and matplotlib around them. The deep learning parts add TensorFlow. Everything installs on Windows, macOS or Linux, or runs in Google Colab without an install.

Installing NLTK · from the Complete NLP Machine Learning in One Shot video · 31:44 to 34:02

The video opens the NLTK site ("NLTK is a leading platform for building Python programs to work with human language data"), mentions spaCy as the other open-source library, and installs NLTK from a notebook cell with !pip install nltk. It sets one task: find the differences between NLTK and spaCy. The table near the end of this lesson answers it.

The clip installs NLTK 3.7 on Windows. Today the same line installs NLTK 3.10, and its tokenizers need language data downloaded once, which the clip's machine already had. Both steps are below.

Checking your Python version

The stack needs Python 3.12 or 3.13: NumPy 2.5 does not install on older versions, and TensorFlow, for the deep learning parts, does not yet publish builds for newer ones. Open a terminal (on Windows, PowerShell; on macOS, the Terminal app) and ask Python for its version:

python3 --version

If it prints 3.12 or 3.13, you are set. Otherwise install Python 3.12 from python.org (on Windows, tick Add python.exe to PATH), or take the uv route below, which downloads Python 3.12 for the project by itself.

Installing the libraries

Pick one route. On your own computer prefer uv: it makes a project folder with its own environment, so these pinned versions never clash with other Python work. pip into your main Python is quicker to type. Colab needs no install at all. The libraries are pinned so your outputs match the lessons; JupyterLab is unpinned.

curl -LsSf https://astral.sh/uv/install.sh | sh
uv init nlp-course --python 3.12
cd nlp-course
uv add nltk==3.10.3 scikit-learn==1.9.1 gensim==4.4.0 numpy==2.5.3 pandas==3.0.6 matplotlib==3.11.2 jupyterlab

What the uv route sets up

  • The first line installs uv, a fast Python package manager. Open a new terminal afterwards so the uv command is found.
  • uv init nlp-course --python 3.12 creates the folder nlp-course with a pyproject.toml that lists the project's libraries, and pins the project to Python 3.12, downloading it if your computer does not have it.
  • uv add creates an environment in nlp-course/.venv, installs the libraries into it and records the exact versions in uv.lock.
  • Run code with uv run, for example uv run python check_setup.py. There is no activation step.

What the pip and Colab routes do

  • python3 -m pip (or py -m pip on Windows) installs into the same Python that python3 (or py) runs, which avoids installing into one Python and running another.
  • Colab already has NumPy, pandas and matplotlib, so its line installs only the three NLP libraries. Restart the session afterwards (Runtime, Restart session) so the new versions load.

Downloading the NLTK data

pip install nltk installs NLTK's code but not its data: the trained sentence tokenizer, the stopword lists, the WordNet dictionary, the tagger and chunker models. Each is downloaded once with nltk.download and saved in a folder called nltk_data, in your home folder by default (on Windows under AppData\Roaming, the path that appears in the videos' notebooks). Download all of them for the course in one go:

python
import nltk

for resource in ["punkt_tab", "stopwords", "wordnet", "omw-1.4",
                 "averaged_perceptron_tagger_eng", "maxent_ne_chunker_tab", "words"]:
    nltk.download(resource)    # prints a log line and True for each

The same download works from a terminal: python -m nltk.downloader punkt_tab stopwords wordnet omw-1.4 averaged_perceptron_tagger_eng maxent_ne_chunker_tab words. What each one is for:

ResourceWhat it isUsed by
punkt_tabPunkt, a trained model of where English sentences endsent_tokenize, word_tokenize
stopwordsStopword lists for 33 languagesstopwords.words('english')
wordnetThe WordNet dictionary of English lemmas, about 11 MBWordNetLemmatizer
omw-1.4The Open Multilingual Wordnet, WordNet in other languagesWordNet outside English; some NLTK versions ask for it with wordnet
averaged_perceptron_tagger_engThe trained English part-of-speech taggerpos_tag
maxent_ne_chunker_tabThe trained named entity chunkerne_chunk
wordsA list of English words that the chunker checks names againstne_chunk

Writing code in Jupyter or VS Code

The videos write their practicals in Jupyter notebooks: code in cells, each cell's output under it. Start JupyterLab from the project folder and it opens in your browser:

uv run jupyter lab

Choose File, New, Notebook, pick the Python 3 kernel, and paste a lesson's code into a cell. In VS Code, install the Python and Jupyter extensions, open the nlp-course folder and run Python: Select Interpreter from the Command Palette: choose the one in .venv for the uv route.

Checking the installed versions

Save this as check_setup.py and run it with uv run python check_setup.py, python3 check_setup.py or py check_setup.py, or paste it into a notebook cell:

ExampleRun on the course's own install
import sys
import nltk, sklearn, gensim, numpy, pandas, matplotlib

print("Python      ", sys.version.split()[0])
print("nltk        ", nltk.__version__)
print("scikit-learn", sklearn.__version__)
print("gensim      ", gensim.__version__)
print("numpy       ", numpy.__version__)
print("pandas      ", pandas.__version__)
print("matplotlib  ", matplotlib.__version__)

Checking every NLTK resource

A version number does not prove the data is there. This check calls each tool once; a missing resource raises LookupError, and the check prints the download line that fixes it:

ExampleRun on NLTK 3.10.3
import nltk

checks = {
    "punkt_tab": lambda: nltk.word_tokenize("Hello world."),
    "stopwords": lambda: len(nltk.corpus.stopwords.words("english")),
    "wordnet": lambda: nltk.stem.WordNetLemmatizer().lemmatize("going", pos="v"),
    "averaged_perceptron_tagger_eng": lambda: nltk.pos_tag(["NLP", "is", "fun"]),
    "maxent_ne_chunker_tab, words": lambda: nltk.ne_chunk(nltk.pos_tag(["Paris", "is", "in", "France"])).pformat(margin=200),
}
for name, check in checks.items():
    try:
        print(f"{name:31} ok  {check()}")
    except LookupError as err:
        missing = str(err).split("'")[1]          # the resource named in the message
        print(f"{name:31} missing: nltk.download('{missing}')")

Reading the checks

  • Python 3.12.13 and nltk 3.10.3 are the versions every output in these lessons comes from; the other four lines match the pins in the install line.
  • punkt_tab split "Hello world." into three tokens, the full stop being one of them.
  • 198 is the size of NLTK's English stopword list today.
  • wordnet turned "going" into "go", read as a verb.
  • The tagger and the chunker labelled the words and found Paris and France as places (GPE). A line that says missing names the exact download to run.

Installing TensorFlow for the deep learning parts

Parts 6 and 7 build embedding layers and LSTMs in Keras, inside TensorFlow. Their practicals show the videos' Colab notebooks with the output they printed, so you can read them without TensorFlow. To run them yourself, add TensorFlow to the same project, or use Colab, which already has it:

uv add tensorflow==2.21.0

This installs the CPU build, which runs every lesson. The Day 9 notebook starts with !pip install tensorflow-gpu; that package no longer exists, and tensorflow is the one to install everywhere. For an NVIDIA GPU on Linux or WSL2, or the GPU of an Apple Silicon Mac, see Installing TensorFlow in the Deep Learning course; in Colab, choose Runtime, Change runtime type, T4 GPU.

Fixing a failed install

Most problems print one of a few messages. Find yours:

  • LookupError: Resource 'punkt_tab' not found. followed by >>> nltk.download('punkt_tab'): the data is missing. Run the line it prints. The same message names stopwords, wordnet, averaged_perceptron_tagger_eng or words for the other tools.
  • The download worked but the error stays, with the old name punkt or averaged_perceptron_tagger downloaded: NLTK 3.10 loads the newer resources punkt_tab and averaged_perceptron_tagger_eng. Download those.
  • [SSL: CERTIFICATE_VERIFY_FAILED] from nltk.download on macOS: the Python from python.org has no root certificates yet. Run Install Certificates.command in its folder under Applications, then download again.
  • No matching distribution found for numpy==2.5.3: the Python running pip is older than 3.12. Check the version, then use a 3.12 install or uv init --python 3.12.
  • error: externally-managed-environment: this Python belongs to the operating system or Homebrew, which blocks pip from changing it. Use the uv route.
  • ModuleNotFoundError: No module named 'nltk': the code runs with a different Python from the one you installed into. Run it with uv run, or pick the .venv interpreter in VS Code.
  • A fresh Colab session raises LookupError again: Colab forgets downloads when the session ends. Keep the download cell at the top of the notebook.

NLTK vs spaCy

The video's question, answered. Both are open-source Python libraries for NLP; they are built for different jobs:

NLTKspaCy
Built forTeaching and research: many algorithms side by sideProduction pipelines: one tuned algorithm per task
How you use itOne function per step: word_tokenize, pos_tag, ne_chunkOne call, nlp(text), returns a document with tokens, lemmas, tags and entities
ModelsSmall classic models and word lists, downloaded with nltk.downloadTrained pipelines per language, such as en_core_web_sm
Choice of algorithmSeveral stemmers, tokenizers and corpora to compareFixed: no stemmer, a lemmatizer built in
In this courseEvery preprocessing lessonShown next to NLTK for named entities

uv vs pip vs Colab

uvpipColab
Where it runsYour computerYour computerA browser, on Google's machines
Keeps versions per projectYes, in pyproject.toml and uv.lockNo, one shared PythonPer session; reinstall after a reset
Installs Python for youYesNoPython is already there
NLTK dataDownloaded once, keptDownloaded once, keptDownloaded again every session

Where you use each setup

  • Following the lessons: any route runs every NLTK, scikit-learn and gensim example.
  • Your own project: the uv route records the exact versions, so a teammate gets the same install with uv sync.
  • A computer you cannot install on, or the Keras practicals without a local TensorFlow: Colab.
Watch out. A program that runs on your machine can fail on a server with LookupError, because nltk_data lives in your home folder and not in the project. Download the resources in the server's setup step, or set the NLTK_DATA environment variable to a folder you ship with the project.
Try it yourself
  • Run import nltk; print(nltk.data.path) to see every folder NLTK searches for its data, in order.
  • Print nltk.corpus.stopwords.fileids() to list the 33 stopword languages, and check whether Hindi is one of them.
  • In a new Colab notebook, run the resource check before any download and read which line fails first.

Little by little, you're building something great.