Installing Python for NLP
NLTK (Natural Language Toolkit) is a Python library of text-processing tools and language data: tokenizers, stemmers, a lemmatizer, stopword lists, a part-of-speech tagger and a named entity chunker.
Last updated: 07 Oct, 2026 · NLTK 3.10 · scikit-learn 1.9
The machine learning parts of the course run on a small stack: NLTK for preprocessing, scikit-learn for vectors and classifiers, gensim for Word2Vec, and NumPy, pandas and matplotlib around them. The deep learning parts add TensorFlow. Everything installs on Windows, macOS or Linux, or runs in Google Colab without an install.
The video opens the NLTK site ("NLTK is a leading platform for building Python programs to work with human language data"), mentions spaCy as the other open-source library, and installs NLTK from a notebook cell with !pip install nltk. It sets one task: find the differences between NLTK and spaCy. The table near the end of this lesson answers it.
The clip installs NLTK 3.7 on Windows. Today the same line installs NLTK 3.10, and its tokenizers need language data downloaded once, which the clip's machine already had. Both steps are below.
Checking your Python version
The stack needs Python 3.12 or 3.13: NumPy 2.5 does not install on older versions, and TensorFlow, for the deep learning parts, does not yet publish builds for newer ones. Open a terminal (on Windows, PowerShell; on macOS, the Terminal app) and ask Python for its version:
python3 --versionIf it prints 3.12 or 3.13, you are set. Otherwise install Python 3.12 from python.org (on Windows, tick Add python.exe to PATH), or take the uv route below, which downloads Python 3.12 for the project by itself.
Installing the libraries
Pick one route. On your own computer prefer uv: it makes a project folder with its own environment, so these pinned versions never clash with other Python work. pip into your main Python is quicker to type. Colab needs no install at all. The libraries are pinned so your outputs match the lessons; JupyterLab is unpinned.
curl -LsSf https://astral.sh/uv/install.sh | sh
uv init nlp-course --python 3.12
cd nlp-course
uv add nltk==3.10.3 scikit-learn==1.9.1 gensim==4.4.0 numpy==2.5.3 pandas==3.0.6 matplotlib==3.11.2 jupyterlabWhat the uv route sets up
- The first line installs uv, a fast Python package manager. Open a new terminal afterwards so the
uvcommand is found. uv init nlp-course --python 3.12creates the foldernlp-coursewith apyproject.tomlthat lists the project's libraries, and pins the project to Python 3.12, downloading it if your computer does not have it.uv addcreates an environment innlp-course/.venv, installs the libraries into it and records the exact versions inuv.lock.- Run code with
uv run, for exampleuv run python check_setup.py. There is no activation step.
What the pip and Colab routes do
python3 -m pip(orpy -m pipon Windows) installs into the same Python thatpython3(orpy) runs, which avoids installing into one Python and running another.- Colab already has NumPy, pandas and matplotlib, so its line installs only the three NLP libraries. Restart the session afterwards (Runtime, Restart session) so the new versions load.
Downloading the NLTK data
pip install nltk installs NLTK's code but not its data: the trained sentence tokenizer, the stopword lists, the WordNet dictionary, the tagger and chunker models. Each is downloaded once with nltk.download and saved in a folder called nltk_data, in your home folder by default (on Windows under AppData\Roaming, the path that appears in the videos' notebooks). Download all of them for the course in one go:
import nltk
for resource in ["punkt_tab", "stopwords", "wordnet", "omw-1.4",
"averaged_perceptron_tagger_eng", "maxent_ne_chunker_tab", "words"]:
nltk.download(resource) # prints a log line and True for eachThe same download works from a terminal: python -m nltk.downloader punkt_tab stopwords wordnet omw-1.4 averaged_perceptron_tagger_eng maxent_ne_chunker_tab words. What each one is for:
| Resource | What it is | Used by |
|---|---|---|
punkt_tab | Punkt, a trained model of where English sentences end | sent_tokenize, word_tokenize |
stopwords | Stopword lists for 33 languages | stopwords.words('english') |
wordnet | The WordNet dictionary of English lemmas, about 11 MB | WordNetLemmatizer |
omw-1.4 | The Open Multilingual Wordnet, WordNet in other languages | WordNet outside English; some NLTK versions ask for it with wordnet |
averaged_perceptron_tagger_eng | The trained English part-of-speech tagger | pos_tag |
maxent_ne_chunker_tab | The trained named entity chunker | ne_chunk |
words | A list of English words that the chunker checks names against | ne_chunk |
Writing code in Jupyter or VS Code
The videos write their practicals in Jupyter notebooks: code in cells, each cell's output under it. Start JupyterLab from the project folder and it opens in your browser:
uv run jupyter labChoose File, New, Notebook, pick the Python 3 kernel, and paste a lesson's code into a cell. In VS Code, install the Python and Jupyter extensions, open the nlp-course folder and run Python: Select Interpreter from the Command Palette: choose the one in .venv for the uv route.
Checking the installed versions
Save this as check_setup.py and run it with uv run python check_setup.py, python3 check_setup.py or py check_setup.py, or paste it into a notebook cell:
import sys
import nltk, sklearn, gensim, numpy, pandas, matplotlib
print("Python ", sys.version.split()[0])
print("nltk ", nltk.__version__)
print("scikit-learn", sklearn.__version__)
print("gensim ", gensim.__version__)
print("numpy ", numpy.__version__)
print("pandas ", pandas.__version__)
print("matplotlib ", matplotlib.__version__)Python 3.12.13 nltk 3.10.3 scikit-learn 1.9.1 gensim 4.4.0 numpy 2.5.3 pandas 3.0.6 matplotlib 3.11.2
Checking every NLTK resource
A version number does not prove the data is there. This check calls each tool once; a missing resource raises LookupError, and the check prints the download line that fixes it:
import nltk
checks = {
"punkt_tab": lambda: nltk.word_tokenize("Hello world."),
"stopwords": lambda: len(nltk.corpus.stopwords.words("english")),
"wordnet": lambda: nltk.stem.WordNetLemmatizer().lemmatize("going", pos="v"),
"averaged_perceptron_tagger_eng": lambda: nltk.pos_tag(["NLP", "is", "fun"]),
"maxent_ne_chunker_tab, words": lambda: nltk.ne_chunk(nltk.pos_tag(["Paris", "is", "in", "France"])).pformat(margin=200),
}
for name, check in checks.items():
try:
print(f"{name:31} ok {check()}")
except LookupError as err:
missing = str(err).split("'")[1] # the resource named in the message
print(f"{name:31} missing: nltk.download('{missing}')")punkt_tab ok ['Hello', 'world', '.']
stopwords ok 198
wordnet ok go
averaged_perceptron_tagger_eng ok [('NLP', 'NNP'), ('is', 'VBZ'), ('fun', 'NN')]
maxent_ne_chunker_tab, words ok (S (GPE Paris/NNP) is/VBZ in/IN (GPE France/NNP))Reading the checks
- Python 3.12.13 and nltk 3.10.3 are the versions every output in these lessons comes from; the other four lines match the pins in the install line.
- punkt_tab split "Hello world." into three tokens, the full stop being one of them.
- 198 is the size of NLTK's English stopword list today.
- wordnet turned "going" into "go", read as a verb.
- The tagger and the chunker labelled the words and found Paris and France as places (GPE). A line that says
missingnames the exact download to run.
Installing TensorFlow for the deep learning parts
Parts 6 and 7 build embedding layers and LSTMs in Keras, inside TensorFlow. Their practicals show the videos' Colab notebooks with the output they printed, so you can read them without TensorFlow. To run them yourself, add TensorFlow to the same project, or use Colab, which already has it:
uv add tensorflow==2.21.0This installs the CPU build, which runs every lesson. The Day 9 notebook starts with !pip install tensorflow-gpu; that package no longer exists, and tensorflow is the one to install everywhere. For an NVIDIA GPU on Linux or WSL2, or the GPU of an Apple Silicon Mac, see Installing TensorFlow in the Deep Learning course; in Colab, choose Runtime, Change runtime type, T4 GPU.
Fixing a failed install
Most problems print one of a few messages. Find yours:
LookupError: Resource 'punkt_tab' not found.followed by>>> nltk.download('punkt_tab'): the data is missing. Run the line it prints. The same message namesstopwords,wordnet,averaged_perceptron_tagger_engorwordsfor the other tools.- The download worked but the error stays, with the old name
punktoraveraged_perceptron_taggerdownloaded: NLTK 3.10 loads the newer resourcespunkt_tabandaveraged_perceptron_tagger_eng. Download those. [SSL: CERTIFICATE_VERIFY_FAILED]fromnltk.downloadon macOS: the Python from python.org has no root certificates yet. Run Install Certificates.command in its folder under Applications, then download again.No matching distribution found for numpy==2.5.3: the Python running pip is older than 3.12. Check the version, then use a 3.12 install oruv init --python 3.12.error: externally-managed-environment: this Python belongs to the operating system or Homebrew, which blocks pip from changing it. Use the uv route.ModuleNotFoundError: No module named 'nltk': the code runs with a different Python from the one you installed into. Run it withuv run, or pick the.venvinterpreter in VS Code.- A fresh Colab session raises
LookupErroragain: Colab forgets downloads when the session ends. Keep the download cell at the top of the notebook.
NLTK vs spaCy
The video's question, answered. Both are open-source Python libraries for NLP; they are built for different jobs:
| NLTK | spaCy | |
|---|---|---|
| Built for | Teaching and research: many algorithms side by side | Production pipelines: one tuned algorithm per task |
| How you use it | One function per step: word_tokenize, pos_tag, ne_chunk | One call, nlp(text), returns a document with tokens, lemmas, tags and entities |
| Models | Small classic models and word lists, downloaded with nltk.download | Trained pipelines per language, such as en_core_web_sm |
| Choice of algorithm | Several stemmers, tokenizers and corpora to compare | Fixed: no stemmer, a lemmatizer built in |
| In this course | Every preprocessing lesson | Shown next to NLTK for named entities |
uv vs pip vs Colab
| uv | pip | Colab | |
|---|---|---|---|
| Where it runs | Your computer | Your computer | A browser, on Google's machines |
| Keeps versions per project | Yes, in pyproject.toml and uv.lock | No, one shared Python | Per session; reinstall after a reset |
| Installs Python for you | Yes | No | Python is already there |
| NLTK data | Downloaded once, kept | Downloaded once, kept | Downloaded again every session |
Where you use each setup
- Following the lessons: any route runs every NLTK, scikit-learn and gensim example.
- Your own project: the uv route records the exact versions, so a teammate gets the same install with
uv sync. - A computer you cannot install on, or the Keras practicals without a local TensorFlow: Colab.
LookupError, because nltk_data lives in your home folder and not in the project. Download the resources in the server's setup step, or set the NLTK_DATA environment variable to a folder you ship with the project.Related
- Previous: Natural language processing (NLP)
- Next: NLP use cases
- Reference: Installing NLTK, Installing NLTK data and the uv documentation
- Run
import nltk; print(nltk.data.path)to see every folder NLTK searches for its data, in order. - Print
nltk.corpus.stopwords.fileids()to list the 33 stopword languages, and check whether Hindi is one of them. - In a new Colab notebook, run the resource check before any download and read which line fails first.
Little by little, you're building something great.