Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Bidirectional LSTM

A bidirectional LSTM is a recurrent layer made of two LSTMs, one reading the sequence left to right and one right to left, whose states are joined at every step so that each position sees the words before it and the words after it.

Last updated: 07 Oct, 2026 · TensorFlow 2 / Keras · NumPy

An LSTM reads in one direction, so at any word it knows only the words that came before. The LSTM text classification (fake news) lesson read whole titles; this lesson adds the second direction and runs the same classifier with it.

Why the words after a blank matter · from Day 11 of the Live NLP series · 11:55 to 14:36

Using the words after the blank

The video's example sentence is "Krish likes to eat ___ in Bangalore": the word to predict sits in the middle. In an ordinary RNN or LSTM, the output y13 at the blank has the context of x11, x12 and x13, the words up to it, because those were fed in before.

It has no context from the words after it, and those decide the answer. Followed by "Bangalore", the blank might be shawarma; change the city to Lucknow in Uttar Pradesh and the likely answer becomes Tunde kebab. Filling a blank in the middle of a sentence needs both sides.

Adding a backward layer · from Day 11 of the Live NLP series · 14:56 to 18:58

Adding a backward layer

The fix is a second network of the same kind, drawn in another colour, connected in the reverse direction: it reads x14 first and passes its state backwards, so information from the later words reaches y13. At every step the outputs of both networks are combined. That pair is a bidirectional RNN, and with LSTM cells a bidirectional LSTM; trained, it can use both sides of the sentence.

A bidirectional LSTM over the sentence Krish likes to eat blank in Lucknow: an orange forward row reads the words left to right and a blue backward row reads them right to left, each with its own weights and a zero initial state; at the blank the two states are concatenated, so the prediction uses the words before it and the word Lucknow after it.
A bidirectional layer: two LSTMs and a joined output
  • The two directions have separate weights and each starts from a zero state, the forward one before the first word and the backward one after the last.
  • Keras joins them by concatenation by default (merge_mode="concat"): 100 forward units and 100 backward units give 200 numbers per step. "sum", "mul" and "ave" are the other options.
  • The output layer is the usual one: a sigmoid or softmax for classification, a linear unit for regression.

Running a bidirectional pass in NumPy

The code builds two small tanh RNNs, runs one forwards and one over the reversed sentence, and joins their states. Then it changes only the last word and checks which half of the blank's state moves.

One direction of the pass

python
def run(X, W_x, W_h):
    h, out = np.zeros(H), []
    for x in X:
        h = np.tanh(W_x @ x + W_h @ h)
        out.append(h)
    return np.array(out)

Joining the two directions

python
def bidirectional(X):
    fwd = run(X, Wf_x, Wf_h)                  # reads left to right
    bwd = run(X[::-1], Wb_x, Wb_h)[::-1]      # reads right to left, put back in order
    return np.concatenate([fwd, bwd], axis=1) # [h_forward ; h_backward] at every step
ExampleRun with NumPy
import numpy as np

rng = np.random.default_rng(42)
words = ["likes", "to", "eat", "___", "in", "Lucknow"]
D, H = 4, 2
X = rng.normal(0, 1, (len(words), D))
Wf_x, Wf_h = rng.normal(0, 0.5, (H, D)), rng.normal(0, 0.5, (H, H))   # forward RNN
Wb_x, Wb_h = rng.normal(0, 0.5, (H, D)), rng.normal(0, 0.5, (H, H))   # backward RNN, its own weights

def run(X, W_x, W_h):
    h, out = np.zeros(H), []
    for x in X:
        h = np.tanh(W_x @ x + W_h @ h)
        out.append(h)
    return np.array(out)


def bidirectional(X):
    fwd = run(X, Wf_x, Wf_h)                  # reads left to right
    bwd = run(X[::-1], Wb_x, Wb_h)[::-1]      # reads right to left, put back in order
    return np.concatenate([fwd, bwd], axis=1) # [h_forward ; h_backward] at every step


out = bidirectional(X)
blank = words.index("___")
print("output shape:", out.shape)
print("at the blank: forward", np.round(out[blank, :H], 4), " backward", np.round(out[blank, H:], 4))

X2 = X.copy()
X2[-1] = rng.normal(0, 1, D)                  # change the last word ("Lucknow" -> another city)
out2 = bidirectional(X2)
print("forward part changed: ", not np.allclose(out[blank, :H], out2[blank, :H]))
print("backward part changed:", not np.allclose(out[blank, H:], out2[blank, H:]))
lstm = 4 * (100 * (100 + 40) + 100)
print("Keras Bidirectional(LSTM(100)) on 40 numbers:", 2 * lstm, " output size:", 2 * 100,
      " model total:", 5000 * 40 + 2 * lstm + 200 + 1)

What the two directions show

  • The output is 6 × 4: 6 words, 2 forward numbers and 2 backward numbers each.
  • Changing "Lucknow" leaves the forward half at the blank unchanged: the forward pass has not reached the last word yet when it is at the blank.
  • The backward half changes, because the backward pass starts at the last word. This is the context the blank was missing.
  • Keras's Bidirectional(LSTM(100)) on 40-number embeddings has 112,800 recurrent parameters, twice the LSTM's 56,400, and outputs 200 numbers; with the Embedding and a Dense unit on top the model has 313,001.

Building the bidirectional fake news classifier

The notebook repeats the fake news classifier with one change in the model. It reads train.csv with the Python parser and skips a broken line:

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
## Dataset: https://www.kaggle.com/competitions/fake-news/data?select=train.csv
import pandas as pd
df=pd.read_csv('train.csv',engine='python',error_bad_lines=False)

pandas 2.0 removed error_bad_lines; today the same call is pd.read_csv('train.csv', engine='python', on_bad_lines='skip').

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
df.isnull().sum()
ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
df.shape

This copy has 7,098 rows, the first third of the 20,800-row file: the upload to Colab had not finished, and the records after row 7,098 never arrived. The missing values, 195 titles, 677 authors and 14 texts, match those first 7,098 records exactly. After dropna 6,226 rows are left:

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
df=df.dropna()
## get the independent features
X=df.drop('label',axis=1)
y=df['label']
## Check whether dataset is balanced or not
y.value_counts()

pandas 3.0 prints the same counts with the index named label and Name: count. The cleaning, one_hot with 5,000 buckets and padding to 20 are those of the LSTM text classification (fake news) lesson, with reset_index run before the loop and padding='pre'. The model wraps the LSTM in Bidirectional:

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
## Creating model
embedding_vector_features=40
model=Sequential()
model.add(Embedding(voc_size,embedding_vector_features,input_length=sent_length))
model.add(Bidirectional(LSTM(100)))
model.add(Dense(1,activation='sigmoid'))
model.compile(loss='binary_crossentropy',optimizer='adam',metrics=['accuracy'])
print(model.summary())

The Bidirectional layer outputs (None, 200), the concatenated last states of the two directions, and holds 112,800 parameters, twice the 56,400 of one LSTM(100). The Dense layer reads 200 numbers, so it has 201 parameters. Training uses batches of 32:

python
import numpy as np
X_final=np.array(embedded_docs)
y_final=np.array(y)
## train test split

from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X_final, y_final, test_size=0.33, random_state=42)
ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
## Model Training
model.fit(X_train,y_train,validation_data=(X_test,y_test),epochs=10,batch_size=32)
python
y_pred=model.predict(X_test)
import numpy as np

y_pred=np.where(y_pred>=0.5,1,0)

The video also tries model.predict_classes(X_test), which Keras removed; for a sigmoid output the replacement is (model.predict(X_test) > 0.5).astype("int32"), which is what the np.where line does.

ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
from sklearn.metrics import confusion_matrix,accuracy_score
confusion_matrix(y_test,y_pred)
ExampleOutput from the video's notebook (TensorFlow 2.8, Colab)
print(accuracy_score(y_test,y_pred))
ExampleRun with matplotlib, on the saved log
import numpy as np
import matplotlib.pyplot as plt

# copied from the saved training log above
loss = [0.3781, 0.1577, 0.0754, 0.0402, 0.0172, 0.0043, 0.0029, 0.0015, 0.0012, 0.0011]
val_loss = [0.2300, 0.2167, 0.2756, 0.3096, 0.4017, 0.5232, 0.5790, 0.6424, 0.6609, 0.6765]
val_acc = [0.8983, 0.9109, 0.9100, 0.9061, 0.9085, 0.9002, 0.9022, 0.9027, 0.9012, 0.9012]
epochs = np.arange(1, 11)
best = int(np.argmin(val_loss)) + 1
print("lowest val_loss:", min(val_loss), "at epoch", best, "with val_accuracy", val_acc[best - 1])
print("epoch 10: val_loss", val_loss[-1], "val_accuracy", val_acc[-1])

plt.figure(figsize=(7, 3.8))
plt.plot(epochs, loss, marker="o", label="training loss")
plt.plot(epochs, val_loss, marker="o", label="validation loss")
plt.axvline(best, color="grey", linestyle="--")
plt.xlabel("epoch")
plt.ylabel("binary cross-entropy")
plt.title("Bidirectional LSTM on fake news titles: training vs validation loss")
plt.legend()
plt.show()
Training loss of the bidirectional LSTM falls from 0.38 to 0.001 over 10 epochs while validation loss is lowest at epoch 2, 0.2167, and rises to 0.6765 by epoch 10, with a dashed line at epoch 2.

Reading the bidirectional run

  • 131 steps per epoch: 4,171 training rows in batches of 32.
  • Training accuracy reaches 0.9998; the 0.9012 accuracy is on the 2,055 test titles (1,077 + 104 + 99 + 775).
  • Validation loss is lowest at epoch 2 (0.2167, accuracy 0.9109) and triples by epoch 10: the same overfitting as the one-direction model, faster.
  • 0.9012 cannot be compared with the LSTM's 0.9037. This run used 6,226 rows, the other 18,285; a fair comparison trains both on the same rows and the same split.

Returning one state or one per word

Bidirectional(LSTM(100)) returns only the final state of each direction, joined: the forward state after the last word and the backward state after it has read back to the first word. That is a many-to-one layer for classifying a whole title. To label every word, ask for every step:

python
from tensorflow.keras.layers import Bidirectional, LSTM

Bidirectional(LSTM(100))                          # output (None, 200): one vector per title
Bidirectional(LSTM(100, return_sequences=True))   # output (None, 20, 200): one vector per word

Bidirectional LSTM vs LSTM

LSTMBidirectional LSTM
Readsleft to rightboth directions, with separate weights
Context at a wordthe words before itthe words before and after it
Recurrent parameters, 100 units on 40 inputs56,400112,800
Output size, 100 units100200 (concat)
Needs the whole sequence firstno, it can run as words arriveyes
Next-word predictionyesno: the backward pass would see the answer

Where you use a bidirectional LSTM

  • Tagging every word: named entity recognition and parts-of-speech tagging, with return_sequences=True.
  • Classifying a complete text, such as a title or a review, where the whole input is available.
  • Filling a blank in the middle of a sentence, the idea BERT later trained at scale.
Watch out. A bidirectional layer must not be used to predict the next word. Its backward pass starts from the end of the text, so at every position it has already read the word it is asked to predict, and training scores look perfect while the model is useless on new text. Language models that generate text read in one direction only.
Try it yourself
  • In the NumPy example, change the first word instead of the last and check which half of the blank's state changes now.
  • Use np.concatenate's alternative, the sum fwd + bwd, and compare the output shape with the concatenated one.
  • Compute the parameters of Bidirectional(LSTM(50)) on 40 inputs with the formula 2 × 4 × (H × (H + D) + H).

Slow is fine. Stopping is the only problem.