Natural language processing (NLP)
Natural language processing (NLP) is the branch of artificial intelligence that turns human language into numbers a model can learn from, so a program can classify, tag, translate, summarise or generate text.
Last updated: 07 Oct, 2026
Spam filtering is the example the video opens with. An e-mail arrives with the subject "Billionaire" and the body "You won a lottery of Billion $", and the label to predict is 1, spam. A table of numbers goes straight into a model; a sentence does not. NLP is the set of steps that bridges that gap.
Turning text into meaningful vectors
In ordinary machine learning the input features are continuous numbers or categories, and a category can be turned into numbers with one-hot, target or ordinal encoding. When a feature is a whole sentence, such as an e-mail's subject and body, none of those encodings is enough. A model cannot understand human language, and the language may be English today and Chinese tomorrow.
So the text is converted into meaningful vectors: lists of numbers that keep information about what the text says. Whenever the input is text or sentences, this processing is natural language processing, and it is what lets a model solve a use case like spam classification. The same idea sits behind Alexa, Google Home and a voice command such as "switch off the AC".
Climbing the NLP roadmap
The video draws the roadmap as a pyramid and reads it from the bottom up:
- Python is the base: one programming language for every step.
- Step 1, text preprocessing: tokenization (a paragraph into sentences, a sentence into words), stemming, lemmatization and stopwords. This step cleans the input.
- Step 2, text to vectors: bag of words, TF-IDF, unigrams and bigrams. This step converts the input text into vectors, and the vectors should capture the context of the statement. Every later technique, up to transformers and BERT, does the same job better: the better the vectors, the better the use case is solved.
The board continues upward. Word2Vec and average Word2Vec give better vectors than bag of words and TF-IDF. Word2Vec is a shallow neural network with one hidden layer, and average Word2Vec averages its word vectors into one vector per sentence. Above them come the deep learning models: RNN, LSTM RNN and GRU RNN, which read a sentence word by word through an embedding layer, a word-vector table trained with the model, then the transformer and BERT. The bottom steps are machine learning with NLTK or spaCy; the top steps are deep learning with TensorFlow or PyTorch.
Each layer up the pyramid brings a larger model that needs more data and more compute. A larger model is not always the better choice: on a few thousand labelled messages, TF-IDF with a linear model is a strong baseline that a deep network has to beat.
Following the learning path
The course follows the pyramid in eleven parts. The table lists every lesson, part by part.
| Part | Lessons | Starts with |
|---|---|---|
| 1. Getting started | Natural language processing (NLP) · Installing Python for NLP · NLP use cases | Installing Python for NLP |
| 2. Text preprocessing | Tokenization · Tokenization with NLTK · Text cleaning and normalisation · Stemming · Lemmatization · Stopwords · Parts of speech (POS) tagging · Named entity recognition (NER) | Tokenization |
| 3. Turning text into vectors | One-hot encoding for text · Bag of words (BoW) · N-grams · TF-IDF · Cosine similarity for documents | One-hot encoding for text |
| 4. Word embeddings | Word embeddings · Word2Vec · CBOW (continuous bag of words) · Skip-gram · Training Word2Vec with gensim · Average Word2Vec | Word embeddings |
| 5. NLP projects with machine learning | Spam classifier with BoW and TF-IDF · Spam classifier with Average Word2Vec · Sentiment analysis of Kindle reviews | Spam classifier with BoW and TF-IDF |
| 6. Recurrent neural networks | Recurrent neural network (RNN) · Types of RNN (one-to-many, many-to-one, many-to-many) · Backpropagation through time (BPTT) · Vanishing and exploding gradients in RNNs | Recurrent neural network (RNN) |
| 7. LSTM, GRU and their practicals | LSTM (long short-term memory) · GRU (gated recurrent unit) · Embedding layer in Keras · LSTM text classification (fake news) · Bidirectional LSTM | LSTM (long short-term memory) |
| 8. Sequence to sequence and attention | Encoder-decoder (seq2seq) models · Attention mechanism (Bahdanau and Luong) | Encoder-decoder (seq2seq) models |
| 9. Transformer building blocks | Transformers · Transformer architecture · Self-attention · Scaled dot-product attention · Multi-head attention · Position-wise feed-forward network · Positional encoding | Transformers |
| 10. Transformer encoder and decoder | Residual connections and layer normalization · Transformer encoder · Masked self-attention · Transformer decoder · Cross-attention (encoder-decoder attention) · Linear and softmax output layer | Residual connections and layer normalization |
| 11. From transformers to today's models | Subword tokenization (BPE and WordPiece) · BERT, GPT and T5 · Fine-tuning transformers with Hugging Face | Subword tokenization (BPE and WordPiece) |
Learning from three videos
No single video covers NLP from tokenization to transformers, so the path has three stretches, each with its own video. The clips in each lesson come from the video of its stretch:
| Stretch | Video | What it teaches |
|---|---|---|
| Parts 1-5: NLP with machine learning | Complete NLP Machine Learning In One Shot (2023, 3 h 53 min) | The roadmap, tokenization, stemming, lemmatization, stopwords, POS tagging, NER, one-hot encoding, bag of words, TF-IDF, Word2Vec, CBOW, skip-gram and average Word2Vec, on a digital board and in Jupyter notebooks. |
| Parts 6-8: NLP with deep learning | The Live NLP series, Days 6 to 11 (2022, 45 min to 1 h 7 min each) | RNNs and their forward pass, backpropagation through time, the LSTM cell, the Keras embedding layer, an LSTM fake-news classifier and the bidirectional LSTM, with Colab notebooks. |
| Parts 9-11: transformers | Complete Transformers For NLP Deep Learning One Shot (2024, 5 h 1 min) | Why transformers, self-attention with Q, K and V, multi-head attention, positional encoding, layer normalization, the encoder, masked attention, the decoder, cross-attention and the output layer, worked on handwritten notes. |
A few lessons have no video behind them: text cleaning, cosine similarity, training Word2Vec with gensim, GRU, encoder-decoder models, attention before transformers, subword tokenization, BERT and GPT, and fine-tuning. They are built from the videos' notes and the standard references, with the same depth and runnable examples.
Who this course is for
- Learners who know some machine learning and want to work with text: e-mails, reviews, chats, documents.
- Interview preparation. The lessons answer the questions interviews ask, such as "what is the difference between stemming and lemmatization?" and "how does TF-IDF weight words differently from bag of words?".
- Anyone heading for large language models. Tokens, embeddings, attention and the transformer are the parts every modern model is made of.
Preparing what you need
- Python basics: lists, dictionaries, loops, functions and importing a library.
- Machine learning basics for parts 3 to 5, from the Machine Learning course: Train and test split, Naive Bayes, Logistic regression and the Confusion matrix.
- Deep learning basics for parts 6 to 11, from the Deep Learning course: the Perceptron, Backpropagation and weight update, Activation functions, the Vanishing gradient problem and the Adam optimizer.
- A computer with Python 3.12 or 3.13, or Google Colab in a browser. Installing Python for NLP covers both. No API keys and no GPU are needed.
Reading a lesson
Every lesson follows the same order:
- A one-sentence definition, then why the idea exists.
- The video's explanation. A clip of a few minutes sits right above the section it covers, and the text under it makes the same points in the same order.
- The board, redrawn: the diagram and the formula in the video's notation.
- The video's own examples: the Kalam speech, "Taj Mahal is a beautiful Monument", the Eiffel Tower sentence, the "good boy, good girl" sentences, the spam messages, the attention example "The cat sat".
- Code with real output. Small commented pieces, then one example that runs, shown with the output it prints. Keras outputs are the ones saved in the video's own Colab notebooks, and are marked as such.
- A comparison table, where the idea is used, a Watch out note on the common mistake, and two or three small changes to try.
Using the videos' notes
The machine learning video's description links its materials, which serve as the notes for parts 1 to 5: the Complete NLP For ML & Deep Learning folder of The Grand Complete Data Science Materials on GitHub. It holds the board PDF NLP For Machine Learning, the practical notebooks (tokenization, stemming, lemmatization, stopwords, POS tagging, NER, bag of words, TF-IDF, Word2Vec and the spam and Kindle projects) with their saved outputs, the SMS spam and Kindle review datasets, and the RNN and LSTM board PDFs used in parts 6 and 7. The transformer notes are the handwritten PDF in Transformers-Materials. The lessons load the datasets straight from these repositories by URL, so there is nothing to download by hand.
The libraries have moved on since the videos were recorded. Where a clip shows an older call, a line under it names the change, and the lesson's code follows the current release:
| In the videos | Today |
|---|---|
sent_tokenize and word_tokenize with the old punkt model | They load punkt_tab: nltk.download('punkt_tab') |
nltk.download('averaged_perceptron_tagger') | pos_tag loads averaged_perceptron_tagger_eng |
nltk.download('maxent_ne_chunker') | ne_chunk loads maxent_ne_chunker_tab, plus the words list |
nltk.pos_tag("a sentence") tagged single characters | It raises TypeError; pass a list of tokens |
| NLTK's English stopword list had 179 words | It has 198: contractions such as "i'm" and "we've" were added |
pip install tensorflow-gpu and TensorFlow 2.9 in Colab | The package is gone; pip install tensorflow, with Keras 3 inside TensorFlow 2.16 and later |
Related
- Next: Installing Python for NLP
- Then: NLP use cases
- Notes: the NLP materials on GitHub
- Reference: Natural Language Processing with Python, the NLTK book
Every expert started right here.