Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Natural Language Processing logoNatural language processing (NLP)

Natural language processing (NLP) is the branch of artificial intelligence that turns human language into numbers a model can learn from, so a program can classify, tag, translate, summarise or generate text.

Last updated: 07 Oct, 2026

Spam filtering is the example the video opens with. An e-mail arrives with the subject "Billionaire" and the body "You won a lottery of Billion $", and the label to predict is 1, spam. A table of numbers goes straight into a model; a sentence does not. NLP is the set of steps that bridges that gap.

Turning text into meaningful vectors

Why text needs NLP · from the Complete NLP Machine Learning in One Shot video · 2:48 to 6:35

In ordinary machine learning the input features are continuous numbers or categories, and a category can be turned into numbers with one-hot, target or ordinal encoding. When a feature is a whole sentence, such as an e-mail's subject and body, none of those encodings is enough. A model cannot understand human language, and the language may be English today and Chinese tomorrow.

So the text is converted into meaningful vectors: lists of numbers that keep information about what the text says. Whenever the input is text or sentences, this processing is natural language processing, and it is what lets a model solve a use case like spam classification. The same idea sits behind Alexa, Google Home and a voice command such as "switch off the AC".

The e-mail with subject Billionaire and body You won a lottery of Billion dollars is labelled 1 for spam; the text cannot go straight into a model, so NLP turns it into a meaningful vector and a classifier decides spam or ham.

Climbing the NLP roadmap

The NLP roadmap, steps 1 and 2 · from the Complete NLP Machine Learning in One Shot video · 6:49 to 10:55

The video draws the roadmap as a pyramid and reads it from the bottom up:

  • Python is the base: one programming language for every step.
  • Step 1, text preprocessing: tokenization (a paragraph into sentences, a sentence into words), stemming, lemmatization and stopwords. This step cleans the input.
  • Step 2, text to vectors: bag of words, TF-IDF, unigrams and bigrams. This step converts the input text into vectors, and the vectors should capture the context of the statement. Every later technique, up to transformers and BERT, does the same job better: the better the vectors, the better the use case is solved.

The board continues upward. Word2Vec and average Word2Vec give better vectors than bag of words and TF-IDF. Word2Vec is a shallow neural network with one hidden layer, and average Word2Vec averages its word vectors into one vector per sentence. Above them come the deep learning models: RNN, LSTM RNN and GRU RNN, which read a sentence word by word through an embedding layer, a word-vector table trained with the model, then the transformer and BERT. The bottom steps are machine learning with NLTK or spaCy; the top steps are deep learning with TensorFlow or PyTorch.

The NLP roadmap as a pyramid from the bottom: Python, text preprocessing to clean the input, text to vectors with bag of words and TF-IDF, then Word2Vec, all machine learning with NLTK or spaCy; above them deep learning with TensorFlow or PyTorch: an embedding layer with RNN, LSTM and GRU, the transformer and BERT.

Each layer up the pyramid brings a larger model that needs more data and more compute. A larger model is not always the better choice: on a few thousand labelled messages, TF-IDF with a linear model is a strong baseline that a deep network has to beat.

Following the learning path

The course as eleven parts in reading order with every lesson listed; parts one to five follow the NLP machine learning one-shot video, parts six to eight the Live NLP series, and parts nine to eleven the transformers one-shot video.

The course follows the pyramid in eleven parts. The table lists every lesson, part by part.

PartLessonsStarts with
1. Getting startedNatural language processing (NLP) · Installing Python for NLP · NLP use casesInstalling Python for NLP
2. Text preprocessingTokenization · Tokenization with NLTK · Text cleaning and normalisation · Stemming · Lemmatization · Stopwords · Parts of speech (POS) tagging · Named entity recognition (NER)Tokenization
3. Turning text into vectorsOne-hot encoding for text · Bag of words (BoW) · N-grams · TF-IDF · Cosine similarity for documentsOne-hot encoding for text
4. Word embeddingsWord embeddings · Word2Vec · CBOW (continuous bag of words) · Skip-gram · Training Word2Vec with gensim · Average Word2VecWord embeddings
5. NLP projects with machine learningSpam classifier with BoW and TF-IDF · Spam classifier with Average Word2Vec · Sentiment analysis of Kindle reviewsSpam classifier with BoW and TF-IDF
6. Recurrent neural networksRecurrent neural network (RNN) · Types of RNN (one-to-many, many-to-one, many-to-many) · Backpropagation through time (BPTT) · Vanishing and exploding gradients in RNNsRecurrent neural network (RNN)
7. LSTM, GRU and their practicalsLSTM (long short-term memory) · GRU (gated recurrent unit) · Embedding layer in Keras · LSTM text classification (fake news) · Bidirectional LSTMLSTM (long short-term memory)
8. Sequence to sequence and attentionEncoder-decoder (seq2seq) models · Attention mechanism (Bahdanau and Luong)Encoder-decoder (seq2seq) models
9. Transformer building blocksTransformers · Transformer architecture · Self-attention · Scaled dot-product attention · Multi-head attention · Position-wise feed-forward network · Positional encodingTransformers
10. Transformer encoder and decoderResidual connections and layer normalization · Transformer encoder · Masked self-attention · Transformer decoder · Cross-attention (encoder-decoder attention) · Linear and softmax output layerResidual connections and layer normalization
11. From transformers to today's modelsSubword tokenization (BPE and WordPiece) · BERT, GPT and T5 · Fine-tuning transformers with Hugging FaceSubword tokenization (BPE and WordPiece)

Learning from three videos

No single video covers NLP from tokenization to transformers, so the path has three stretches, each with its own video. The clips in each lesson come from the video of its stretch:

StretchVideoWhat it teaches
Parts 1-5: NLP with machine learningComplete NLP Machine Learning In One Shot (2023, 3 h 53 min)The roadmap, tokenization, stemming, lemmatization, stopwords, POS tagging, NER, one-hot encoding, bag of words, TF-IDF, Word2Vec, CBOW, skip-gram and average Word2Vec, on a digital board and in Jupyter notebooks.
Parts 6-8: NLP with deep learningThe Live NLP series, Days 6 to 11 (2022, 45 min to 1 h 7 min each)RNNs and their forward pass, backpropagation through time, the LSTM cell, the Keras embedding layer, an LSTM fake-news classifier and the bidirectional LSTM, with Colab notebooks.
Parts 9-11: transformersComplete Transformers For NLP Deep Learning One Shot (2024, 5 h 1 min)Why transformers, self-attention with Q, K and V, multi-head attention, positional encoding, layer normalization, the encoder, masked attention, the decoder, cross-attention and the output layer, worked on handwritten notes.

A few lessons have no video behind them: text cleaning, cosine similarity, training Word2Vec with gensim, GRU, encoder-decoder models, attention before transformers, subword tokenization, BERT and GPT, and fine-tuning. They are built from the videos' notes and the standard references, with the same depth and runnable examples.

Who this course is for

  • Learners who know some machine learning and want to work with text: e-mails, reviews, chats, documents.
  • Interview preparation. The lessons answer the questions interviews ask, such as "what is the difference between stemming and lemmatization?" and "how does TF-IDF weight words differently from bag of words?".
  • Anyone heading for large language models. Tokens, embeddings, attention and the transformer are the parts every modern model is made of.

Preparing what you need

Reading a lesson

Every lesson follows the same order:

  • A one-sentence definition, then why the idea exists.
  • The video's explanation. A clip of a few minutes sits right above the section it covers, and the text under it makes the same points in the same order.
  • The board, redrawn: the diagram and the formula in the video's notation.
  • The video's own examples: the Kalam speech, "Taj Mahal is a beautiful Monument", the Eiffel Tower sentence, the "good boy, good girl" sentences, the spam messages, the attention example "The cat sat".
  • Code with real output. Small commented pieces, then one example that runs, shown with the output it prints. Keras outputs are the ones saved in the video's own Colab notebooks, and are marked as such.
  • A comparison table, where the idea is used, a Watch out note on the common mistake, and two or three small changes to try.

Using the videos' notes

The machine learning video's description links its materials, which serve as the notes for parts 1 to 5: the Complete NLP For ML & Deep Learning folder of The Grand Complete Data Science Materials on GitHub. It holds the board PDF NLP For Machine Learning, the practical notebooks (tokenization, stemming, lemmatization, stopwords, POS tagging, NER, bag of words, TF-IDF, Word2Vec and the spam and Kindle projects) with their saved outputs, the SMS spam and Kindle review datasets, and the RNN and LSTM board PDFs used in parts 6 and 7. The transformer notes are the handwritten PDF in Transformers-Materials. The lessons load the datasets straight from these repositories by URL, so there is nothing to download by hand.

The libraries have moved on since the videos were recorded. Where a clip shows an older call, a line under it names the change, and the lesson's code follows the current release:

In the videosToday
sent_tokenize and word_tokenize with the old punkt modelThey load punkt_tab: nltk.download('punkt_tab')
nltk.download('averaged_perceptron_tagger')pos_tag loads averaged_perceptron_tagger_eng
nltk.download('maxent_ne_chunker')ne_chunk loads maxent_ne_chunker_tab, plus the words list
nltk.pos_tag("a sentence") tagged single charactersIt raises TypeError; pass a list of tokens
NLTK's English stopword list had 179 wordsIt has 198: contractions such as "i'm" and "we've" were added
pip install tensorflow-gpu and TensorFlow 2.9 in ColabThe package is gone; pip install tensorflow, with Keras 3 inside TensorFlow 2.16 and later
Back toAll courses

Every expert started right here.