Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

LSTM text classification (fake news)

LSTM text classification is classifying a text by reading its words in order with an LSTM layer and predicting a label from the last hidden state; here the texts are news titles and the labels are reliable and unreliable.

Last updated: 07 Oct, 2026 · TensorFlow 2 / Keras · NumPy

The Embedding layer in Keras lesson turned short sentences into vectors. This practical puts the same steps in front of an LSTM (long short-term memory) layer and trains it on the titles of 18,285 news articles.

Getting the fake news dataset

The data is train.csv from the Fake News competition on Kaggle (a Kaggle login is needed to download it). It has 20,800 articles with five columns: id, title, author, text and label, where 1 means unreliable and 0 reliable. The video uploads the file to Colab and predicts the label from the title alone.

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
import pandas as pd
df=pd.read_csv('train.csv')
df.isnull().sum()

A missing title or text cannot be filled in sensibly, so the notebook drops every row with a missing value and splits the columns into features X and the label y:

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
###Drop Nan Values
df=df.dropna()
## Get the Independent Features

X=df.drop('label',axis=1)
## Get the Dependent features
y=df['label']
X.shape
The pipeline of the fake news classifier: train.csv with 20,800 rows, dropna to 18,285 rows, cleaned titles (regex, lower case, stopwords, Porter stemming), one_hot into 5,000 buckets, padding to 20, an Embedding of 5,000 by 40, an LSTM with 100 units, a Dense sigmoid unit, and a threshold on the probability of unreliable.

Preparing the titles

The notebook imports the Keras pieces, sets the vocabulary size to 5,000 and copies X into messages. Because dropna removed rows, the row labels now have gaps; reset_index renumbers them 0, 1, 2 and so on.

python
from tensorflow.keras.layers import Embedding
from tensorflow.keras.preprocessing.sequence import pad_sequences
from tensorflow.keras.models import Sequential
from tensorflow.keras.preprocessing.text import one_hot
from tensorflow.keras.layers import LSTM
from tensorflow.keras.layers import Dense
### Vocabulary size
voc_size=5000
messages=X.copy()
messages.reset_index(inplace=True)

The cleaning loop below reads messages['title'][i], which looks rows up by label, so this reset_index line has to run. Without it the loop stops at the first missing label, here with KeyError: 6. The same failure on a three-row table:

ExampleRun on pandas 3.0
import pandas as pd

df = pd.DataFrame({"title": ["first title", None, "third title"], "label": [1, 0, 1]})
messages = df.dropna()                     # row 1 is dropped, the labels are now 0 and 2
print("index after dropna:", messages.index.tolist())
try:
    for i in range(len(messages)):
        print(messages['title'][i])
except KeyError as err:
    print("KeyError:", err)              # there is no row labelled 1 any more
ExampleRun on pandas 3.0
import pandas as pd

df = pd.DataFrame({"title": ["first title", None, "third title"], "label": [1, 0, 1]})
messages = df.dropna()
messages.reset_index(inplace=True)         # the labels become 0, 1 again
for i in range(len(messages)):
    print(messages['title'][i])
Cleaning the titles with stemming and stopwords · from Day 10 of the Live NLP series · 21:13 to 24:28

Cleaning the titles with stemming and stopwords

The cleaning builds a list called corpus, one cleaned title per row. For each title it replaces everything except the letters a to z and A to Z with a space, lowercases the result, splits it into words, and in a list comprehension keeps the Stemming of every word that is not in the English Stopwords list. The words are joined back into one string. Lemmatization would also work, but on about 18,000 titles it takes longer than stemming.

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
import nltk
import re
from nltk.corpus import stopwords
nltk.download('stopwords')
python
### Dataset Preprocessing
from nltk.stem.porter import PorterStemmer ##stemming purpose
ps = PorterStemmer()
corpus = []
for i in range(0, len(messages)):
    review = re.sub('[^a-zA-Z]', ' ', messages['title'][i])
    review = review.lower()
    review = review.split()
    
    review = [ps.stem(word) for word in review if not word in stopwords.words('english')]
    review = ' '.join(review)
    corpus.append(review)
ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
corpus[1]

The same loop on the first six titles of the file, with today's NLTK. The stopwords become a set built once: calling stopwords.words('english') for every word re-reads the list each time and makes the loop far slower, with the same result.

ExampleRun on NLTK 3.10
import re
from nltk.corpus import stopwords
from nltk.stem.porter import PorterStemmer

titles = [
    'House Dem Aide: We Didn’t Even See Comey’s Letter Until Jason Chaffetz Tweeted It',
    'FLYNN: Hillary Clinton, Big Woman on Campus - Breitbart',
    'Why the Truth Might Get You Fired',
    '15 Civilians Killed In Single US Airstrike Have Been Identified',
    'Iranian woman jailed for fictional unpublished story about woman stoned to death for adultery',
    'Jackie Mason: Hollywood Would Love Trump if He Bombed North Korea over Lack of Trans Bathrooms (Exclusive Video) - Breitbart',
]
ps = PorterStemmer()
stop = set(stopwords.words('english'))      # built once, not once per word
corpus = []
for title in titles:
    review = re.sub('[^a-zA-Z]', ' ', title)
    review = review.lower().split()
    review = [ps.stem(word) for word in review if word not in stop]
    corpus.append(' '.join(review))
for line in corpus:
    print(line)

The six lines match the first six entries of the notebook's saved corpus, word for word. Notice "breitbart" at the end of two of them: the outlet's name is part of the title.

Encoding and padding the titles

one_hot hashes each stemmed word into a number from 1 to 4,999. The second title, 'flynn hillari clinton big woman campu breitbart', becomes seven numbers; in this run "flynn" is 2861:

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
onehot_repr=[one_hot(words,voc_size)for words in corpus] 
onehot_repr[1]

With 13,931 distinct stems and 4,999 buckets, most buckets hold several words, so different words share a number. Every title is then padded to 20 numbers, this time at the end:

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
sent_length=20
embedded_docs=pad_sequences(onehot_repr,padding='post',maxlen=sent_length)
print(embedded_docs)
embedded_docs[1]

A length of 20 covers almost every title, but not all: after cleaning, the longest has 47 words, 37 titles are longer than 20 and lose their first words to the default truncating='pre', and 54 titles clean to nothing and become 20 zeros.

Building the LSTM model · from Day 10 of the Live NLP series · 31:47 to 34:22

Building the LSTM model

The model is a Sequential stack of three layers. The Embedding layer takes the vocabulary size (5,000), the number of features per word (40) and the input length (20), so every index becomes 40 numbers. The LSTM layer has 100 units; 200 or 300 would also work, and the number is a hyperparameter. The output is one Dense unit with a sigmoid, because the label is binary, and a binary output goes with binary cross-entropy as the loss, Adam as the optimizer and accuracy as the metric. The summary reports 256,501 parameters.

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
## Creating model
embedding_vector_features=40 ##features representation
model=Sequential()
model.add(Embedding(voc_size,embedding_vector_features,input_length=sent_length))
model.add(LSTM(100))
model.add(Dense(1,activation='sigmoid'))
model.compile(loss='binary_crossentropy',optimizer='adam',metrics=['accuracy'])
print(model.summary())

On Keras 3, input_length is deprecated and one_hot no longer exists; the section on today's Keras below shows the current form of both.

The 40 numbers per word are learned by backpropagation together with the LSTM, from the binary cross-entropy of this task. The parameter counts follow from the layer sizes:

ExamplePlain Python
voc_size, dim, units = 5000, 40, 100
embedding = voc_size * dim                          # one 40-number row per index
lstm = 4 * (units * (units + dim) + units)          # four layers on [h(t-1), x(t)], one bias each
dense = units * 1 + 1
print("embedding:", embedding, " lstm:", lstm, " dense:", dense, " total:", embedding + lstm + dense)
print("training steps per epoch:", -(-12250 // 64))  # 12,250 training rows in batches of 64

Training and evaluating the model

The padded titles become NumPy arrays and are split, a third for testing (Train and test split):

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
import numpy as np
X_final=np.array(embedded_docs)
y_final=np.array(y)
X_final.shape,y_final.shape
python
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X_final, y_final, test_size=0.33, random_state=42)

The video trains this model for 10 epochs, then adds Dropout with a rate of 0.3 after the Embedding and after the LSTM layer and trains again. The training log saved in the notebook is the run of this second model:

python
from tensorflow.keras.layers import Dropout
## Creating model
embedding_vector_features=40
model=Sequential()
model.add(Embedding(voc_size,embedding_vector_features,input_length=sent_length))
model.add(Dropout(0.3))
model.add(LSTM(100))
model.add(Dropout(0.3))
model.add(Dense(1,activation='sigmoid'))
model.compile(loss='binary_crossentropy',optimizer='adam',metrics=['accuracy'])
ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
### Finally Training
model.fit(X_train,y_train,validation_data=(X_test,y_test),epochs=10,batch_size=64)

The predictions are probabilities; a threshold turns them into 0 or 1. The notebook uses 0.6, then prints the Confusion matrix, the accuracy and the classification report:

python
y_pred=model.predict(X_test)
y_pred=np.where(y_pred > 0.6, 1,0) ##AUC ROC Curve
from sklearn.metrics import confusion_matrix
ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
confusion_matrix(y_test,y_pred)
ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
from sklearn.metrics import accuracy_score
accuracy_score(y_test,y_pred)
ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
from sklearn.metrics import classification_report
print(classification_report(y_test,y_pred))
ExampleRun with matplotlib, on the saved log
import numpy as np
import matplotlib.pyplot as plt

# loss and val_loss copied from the saved training log above (the model with Dropout)
loss = [0.3495, 0.1464, 0.1037, 0.0699, 0.0547, 0.0456, 0.0319, 0.0286, 0.0217, 0.0224]
val_loss = [0.2056, 0.2070, 0.2376, 0.2565, 0.2616, 0.3549, 0.3948, 0.3906, 0.4145, 0.4555]
epochs = np.arange(1, 11)
best = int(np.argmin(val_loss)) + 1
print("lowest val_loss:", min(val_loss), "at epoch", best)
print("val_loss at epoch 10 is", round(val_loss[-1] / min(val_loss), 2), "times the lowest")

plt.figure(figsize=(7, 3.8))
plt.plot(epochs, loss, marker="o", label="training loss")
plt.plot(epochs, val_loss, marker="o", label="validation loss")
plt.axvline(best, color="grey", linestyle="--")
plt.xlabel("epoch")
plt.ylabel("binary cross-entropy")
plt.title("LSTM with dropout on fake news titles: training vs validation loss")
plt.legend()
plt.show()
Training loss falls from 0.35 to 0.02 over 10 epochs while validation loss is lowest at epoch 1, 0.2056, and climbs to 0.4555 by epoch 10, with a dashed line at epoch 1.

Reading the training log and the scores

  • 192 steps per epoch: 12,250 training rows in batches of 64.
  • Training accuracy ends at 0.9918 while validation accuracy stays near 0.90 to 0.92.
  • Validation loss is lowest after epoch 1 (0.2056) and more than doubles by epoch 10: the model overfits after the first pass. Early stopping on val_loss would keep the epoch-1 weights.
  • Test accuracy is 0.9037 on 6,035 titles at a threshold of 0.6, with 284 reliable titles flagged as unreliable and 297 unreliable ones missed.
  • The plain LSTM, before dropout, scored 0.9084 on the same test set at a threshold of 0.5 in the video's run, with the confusion matrix [[3156, 263], [290, 2326]]: dropout did not help here.
  • The validation data is the test set itself. A separate validation split (validation_split=0.1 on the training data) keeps the test score untouched by choices made while training.

Checking what the model learned

Before trusting the score, compare it with a rule that does not read the news at all. The code needs the competition's train.csv:

python
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split

df = pd.read_csv('train.csv').dropna()
title = df['title'].str.lower()
named = title.str.contains('new york times') | title.str.contains('breitbart')
pred = np.where(named, 0, 1)               # names one of two outlets -> reliable, else unreliable
_, p_test, _, y_test = train_test_split(pred, df['label'].to_numpy(), test_size=0.33, random_state=42)
print((p_test == y_test).mean())

It prints 0.8966. Of the 18,285 titles, 46.8% contain "New York Times" or "Breitbart", and only 0.19% of those are labelled unreliable. The LSTM's 0.90 is little more than this rule: most of what it learned is the outlet's name at the end of the title.

Writing the model with today's Keras

python
import keras
from keras import layers

vectorize = layers.TextVectorization(max_tokens=5000, output_sequence_length=20)
vectorize.adapt(corpus)
X_final = vectorize(np.array(corpus))      # a real vocabulary instead of hashing

model = keras.Sequential([
    keras.Input(shape=(20,)),
    layers.Embedding(5000, 40, mask_zero=True),   # skip the padded positions
    layers.LSTM(100),
    layers.Dense(1, activation="sigmoid"),
])
python
model.compile(loss="binary_crossentropy", optimizer="adam", metrics=["accuracy"])
stop = keras.callbacks.EarlyStopping(monitor="val_loss", patience=2, restore_best_weights=True)
model.fit(X_train, y_train, validation_split=0.1, epochs=10, batch_size=64, callbacks=[stop])
  • one_hot is gone in Keras 3; TextVectorization replaces it.
  • input_length is deprecated; keras.Input(shape=(20,)) gives the shape.
  • mask_zero=True stops the LSTM from reading the trailing padding zeros.
  • validation_split and EarlyStopping keep the test set out of training decisions.

Plain LSTM vs LSTM with dropout

Plain LSTMWith Dropout(0.3)
Threshold0.50.6
Test accuracy0.90840.9037
Confusion matrix[[3156, 263], [290, 2326]][[3135, 284], [297, 2319]]
Parameters256,501256,501 (dropout adds none)

Where you use LSTM text classification

  • Moderation and spam filtering, where the order of words changes the meaning.
  • Sentiment and topic labelling of reviews, tickets and posts.
  • A step up from bag-of-words models such as the Spam classifier with BoW and TF-IDF, once those baselines are known.
Watch out. A high score can come from a shortcut in the data. Here a rule that only checks whether the title names the New York Times or Breitbart scores 0.8966 on the same test split, against the LSTM's 0.9037, so the model mostly learned the outlet's name. Strip outlet names from the titles, or test on outlets the model has never seen, before claiming it detects fake news.
Try it yourself
  • Add a seventh title to the cleaning example that ends in " - The New York Times" and look at what its last three stems are.
  • In the curve example, find the epoch where training loss first drops below 0.05 and compare validation loss at that epoch with its lowest value.
  • Change units = 100 to units = 200 in the parameter example and see how much the LSTM layer grows.

You understood something today that you didn't yesterday.