Bidirectional LSTM
A bidirectional LSTM is a recurrent layer made of two LSTMs, one reading the sequence left to right and one right to left, whose states are joined at every step so that each position sees the words before it and the words after it.
Last updated: 07 Oct, 2026 · TensorFlow 2 / Keras · NumPy
An LSTM reads in one direction, so at any word it knows only the words that came before. The LSTM text classification (fake news) lesson read whole titles; this lesson adds the second direction and runs the same classifier with it.
Using the words after the blank
The video's example sentence is "Krish likes to eat ___ in Bangalore": the word to predict sits in the middle. In an ordinary RNN or LSTM, the output y13 at the blank has the context of x11, x12 and x13, the words up to it, because those were fed in before.
It has no context from the words after it, and those decide the answer. Followed by "Bangalore", the blank might be shawarma; change the city to Lucknow in Uttar Pradesh and the likely answer becomes Tunde kebab. Filling a blank in the middle of a sentence needs both sides.
Adding a backward layer
The fix is a second network of the same kind, drawn in another colour, connected in the reverse direction: it reads x14 first and passes its state backwards, so information from the later words reaches y13. At every step the outputs of both networks are combined. That pair is a bidirectional RNN, and with LSTM cells a bidirectional LSTM; trained, it can use both sides of the sentence.
- The two directions have separate weights and each starts from a zero state, the forward one before the first word and the backward one after the last.
- Keras joins them by concatenation by default (
merge_mode="concat"): 100 forward units and 100 backward units give 200 numbers per step."sum","mul"and"ave"are the other options. - The output layer is the usual one: a sigmoid or softmax for classification, a linear unit for regression.
Running a bidirectional pass in NumPy
The code builds two small tanh RNNs, runs one forwards and one over the reversed sentence, and joins their states. Then it changes only the last word and checks which half of the blank's state moves.
One direction of the pass
def run(X, W_x, W_h):
h, out = np.zeros(H), []
for x in X:
h = np.tanh(W_x @ x + W_h @ h)
out.append(h)
return np.array(out)Joining the two directions
def bidirectional(X):
fwd = run(X, Wf_x, Wf_h) # reads left to right
bwd = run(X[::-1], Wb_x, Wb_h)[::-1] # reads right to left, put back in order
return np.concatenate([fwd, bwd], axis=1) # [h_forward ; h_backward] at every stepimport numpy as np
rng = np.random.default_rng(42)
words = ["likes", "to", "eat", "___", "in", "Lucknow"]
D, H = 4, 2
X = rng.normal(0, 1, (len(words), D))
Wf_x, Wf_h = rng.normal(0, 0.5, (H, D)), rng.normal(0, 0.5, (H, H)) # forward RNN
Wb_x, Wb_h = rng.normal(0, 0.5, (H, D)), rng.normal(0, 0.5, (H, H)) # backward RNN, its own weights
def run(X, W_x, W_h):
h, out = np.zeros(H), []
for x in X:
h = np.tanh(W_x @ x + W_h @ h)
out.append(h)
return np.array(out)
def bidirectional(X):
fwd = run(X, Wf_x, Wf_h) # reads left to right
bwd = run(X[::-1], Wb_x, Wb_h)[::-1] # reads right to left, put back in order
return np.concatenate([fwd, bwd], axis=1) # [h_forward ; h_backward] at every step
out = bidirectional(X)
blank = words.index("___")
print("output shape:", out.shape)
print("at the blank: forward", np.round(out[blank, :H], 4), " backward", np.round(out[blank, H:], 4))
X2 = X.copy()
X2[-1] = rng.normal(0, 1, D) # change the last word ("Lucknow" -> another city)
out2 = bidirectional(X2)
print("forward part changed: ", not np.allclose(out[blank, :H], out2[blank, :H]))
print("backward part changed:", not np.allclose(out[blank, H:], out2[blank, H:]))
lstm = 4 * (100 * (100 + 40) + 100)
print("Keras Bidirectional(LSTM(100)) on 40 numbers:", 2 * lstm, " output size:", 2 * 100,
" model total:", 5000 * 40 + 2 * lstm + 200 + 1)output shape: (6, 4) at the blank: forward [-0.5303 0.883 ] backward [-0.7668 -0.0175] forward part changed: False backward part changed: True Keras Bidirectional(LSTM(100)) on 40 numbers: 112800 output size: 200 model total: 313001
What the two directions show
- The output is 6 × 4: 6 words, 2 forward numbers and 2 backward numbers each.
- Changing "Lucknow" leaves the forward half at the blank unchanged: the forward pass has not reached the last word yet when it is at the blank.
- The backward half changes, because the backward pass starts at the last word. This is the context the blank was missing.
- Keras's Bidirectional(LSTM(100)) on 40-number embeddings has 112,800 recurrent parameters, twice the LSTM's 56,400, and outputs 200 numbers; with the Embedding and a Dense unit on top the model has 313,001.
Building the bidirectional fake news classifier
The notebook repeats the fake news classifier with one change in the model. It reads train.csv with the Python parser and skips a broken line:
## Dataset: https://www.kaggle.com/competitions/fake-news/data?select=train.csv
import pandas as pd
df=pd.read_csv('train.csv',engine='python',error_bad_lines=False)Skipping line 7100: unexpected end of data
pandas 2.0 removed error_bad_lines; today the same call is pd.read_csv('train.csv', engine='python', on_bad_lines='skip').
df.isnull().sum()id 0 title 195 author 677 text 14 label 0 dtype: int64
df.shape(7098, 5)
This copy has 7,098 rows, the first third of the 20,800-row file: the upload to Colab had not finished, and the records after row 7,098 never arrived. The missing values, 195 titles, 677 authors and 14 texts, match those first 7,098 records exactly. After dropna 6,226 rows are left:
df=df.dropna()
## get the independent features
X=df.drop('label',axis=1)
y=df['label']
## Check whether dataset is balanced or not
y.value_counts()0 3530 1 2696 Name: label, dtype: int64
pandas 3.0 prints the same counts with the index named label and Name: count. The cleaning, one_hot with 5,000 buckets and padding to 20 are those of the LSTM text classification (fake news) lesson, with reset_index run before the loop and padding='pre'. The model wraps the LSTM in Bidirectional:
## Creating model
embedding_vector_features=40
model=Sequential()
model.add(Embedding(voc_size,embedding_vector_features,input_length=sent_length))
model.add(Bidirectional(LSTM(100)))
model.add(Dense(1,activation='sigmoid'))
model.compile(loss='binary_crossentropy',optimizer='adam',metrics=['accuracy'])
print(model.summary())Model: "sequential"
_________________________________________________________________
Layer (type) Output Shape Param #
=================================================================
embedding (Embedding) (None, 20, 40) 200000
bidirectional (Bidirectiona (None, 200) 112800
l)
dense (Dense) (None, 1) 201
=================================================================
Total params: 313,001
Trainable params: 313,001
Non-trainable params: 0
_________________________________________________________________
NoneThe Bidirectional layer outputs (None, 200), the concatenated last states of the two directions, and holds 112,800 parameters, twice the 56,400 of one LSTM(100). The Dense layer reads 200 numbers, so it has 201 parameters. Training uses batches of 32:
import numpy as np
X_final=np.array(embedded_docs)
y_final=np.array(y)
## train test split
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X_final, y_final, test_size=0.33, random_state=42)## Model Training
model.fit(X_train,y_train,validation_data=(X_test,y_test),epochs=10,batch_size=32)Epoch 1/10 131/131 [==============================] - 15s 22ms/step - loss: 0.3781 - accuracy: 0.8125 - val_loss: 0.2300 - val_accuracy: 0.8983 Epoch 2/10 131/131 [==============================] - 2s 13ms/step - loss: 0.1577 - accuracy: 0.9348 - val_loss: 0.2167 - val_accuracy: 0.9109 Epoch 3/10 131/131 [==============================] - 2s 13ms/step - loss: 0.0754 - accuracy: 0.9751 - val_loss: 0.2756 - val_accuracy: 0.9100 Epoch 4/10 131/131 [==============================] - 2s 14ms/step - loss: 0.0402 - accuracy: 0.9883 - val_loss: 0.3096 - val_accuracy: 0.9061 Epoch 5/10 131/131 [==============================] - 2s 14ms/step - loss: 0.0172 - accuracy: 0.9959 - val_loss: 0.4017 - val_accuracy: 0.9085 Epoch 6/10 131/131 [==============================] - 2s 16ms/step - loss: 0.0043 - accuracy: 0.9993 - val_loss: 0.5232 - val_accuracy: 0.9002 Epoch 7/10 131/131 [==============================] - 2s 13ms/step - loss: 0.0029 - accuracy: 0.9995 - val_loss: 0.5790 - val_accuracy: 0.9022 Epoch 8/10 131/131 [==============================] - 2s 14ms/step - loss: 0.0015 - accuracy: 0.9998 - val_loss: 0.6424 - val_accuracy: 0.9027 Epoch 9/10 131/131 [==============================] - 2s 13ms/step - loss: 0.0012 - accuracy: 0.9998 - val_loss: 0.6609 - val_accuracy: 0.9012 Epoch 10/10 131/131 [==============================] - 2s 13ms/step - loss: 0.0011 - accuracy: 0.9998 - val_loss: 0.6765 - val_accuracy: 0.9012
y_pred=model.predict(X_test)
import numpy as np
y_pred=np.where(y_pred>=0.5,1,0)The video also tries model.predict_classes(X_test), which Keras removed; for a sigmoid output the replacement is (model.predict(X_test) > 0.5).astype("int32"), which is what the np.where line does.
from sklearn.metrics import confusion_matrix,accuracy_score
confusion_matrix(y_test,y_pred)array([[1077, 104],
[ 99, 775]])print(accuracy_score(y_test,y_pred))0.9012165450121654
import numpy as np
import matplotlib.pyplot as plt
# copied from the saved training log above
loss = [0.3781, 0.1577, 0.0754, 0.0402, 0.0172, 0.0043, 0.0029, 0.0015, 0.0012, 0.0011]
val_loss = [0.2300, 0.2167, 0.2756, 0.3096, 0.4017, 0.5232, 0.5790, 0.6424, 0.6609, 0.6765]
val_acc = [0.8983, 0.9109, 0.9100, 0.9061, 0.9085, 0.9002, 0.9022, 0.9027, 0.9012, 0.9012]
epochs = np.arange(1, 11)
best = int(np.argmin(val_loss)) + 1
print("lowest val_loss:", min(val_loss), "at epoch", best, "with val_accuracy", val_acc[best - 1])
print("epoch 10: val_loss", val_loss[-1], "val_accuracy", val_acc[-1])
plt.figure(figsize=(7, 3.8))
plt.plot(epochs, loss, marker="o", label="training loss")
plt.plot(epochs, val_loss, marker="o", label="validation loss")
plt.axvline(best, color="grey", linestyle="--")
plt.xlabel("epoch")
plt.ylabel("binary cross-entropy")
plt.title("Bidirectional LSTM on fake news titles: training vs validation loss")
plt.legend()
plt.show()lowest val_loss: 0.2167 at epoch 2 with val_accuracy 0.9109 epoch 10: val_loss 0.6765 val_accuracy 0.9012
Reading the bidirectional run
- 131 steps per epoch: 4,171 training rows in batches of 32.
- Training accuracy reaches 0.9998; the 0.9012 accuracy is on the 2,055 test titles (1,077 + 104 + 99 + 775).
- Validation loss is lowest at epoch 2 (0.2167, accuracy 0.9109) and triples by epoch 10: the same overfitting as the one-direction model, faster.
- 0.9012 cannot be compared with the LSTM's 0.9037. This run used 6,226 rows, the other 18,285; a fair comparison trains both on the same rows and the same split.
Returning one state or one per word
Bidirectional(LSTM(100)) returns only the final state of each direction, joined: the forward state after the last word and the backward state after it has read back to the first word. That is a many-to-one layer for classifying a whole title. To label every word, ask for every step:
from tensorflow.keras.layers import Bidirectional, LSTM
Bidirectional(LSTM(100)) # output (None, 200): one vector per title
Bidirectional(LSTM(100, return_sequences=True)) # output (None, 20, 200): one vector per wordBidirectional LSTM vs LSTM
| LSTM | Bidirectional LSTM | |
|---|---|---|
| Reads | left to right | both directions, with separate weights |
| Context at a word | the words before it | the words before and after it |
| Recurrent parameters, 100 units on 40 inputs | 56,400 | 112,800 |
| Output size, 100 units | 100 | 200 (concat) |
| Needs the whole sequence first | no, it can run as words arrive | yes |
| Next-word prediction | yes | no: the backward pass would see the answer |
Where you use a bidirectional LSTM
- Tagging every word: named entity recognition and parts-of-speech tagging, with
return_sequences=True. - Classifying a complete text, such as a title or a review, where the whole input is available.
- Filling a blank in the middle of a sentence, the idea BERT later trained at scale.
Related
- Previous: LSTM text classification (fake news)
- Next: Encoder-decoder (seq2seq) models
- See also: Types of RNN (one-to-many, many-to-one, many-to-many)
- Reference: Bidirectional layer in the Keras API
- In the NumPy example, change the first word instead of the last and check which half of the blank's state changes now.
- Use
np.concatenate's alternative, the sumfwd + bwd, and compare the output shape with the concatenated one. - Compute the parameters of
Bidirectional(LSTM(50))on 40 inputs with the formula 2 × 4 × (H × (H + D) + H).
Slow is fine. Stopping is the only problem.