Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

LSTM (long short-term memory)

LSTM (long short-term memory) is a recurrent network cell that keeps a separate cell state and uses three gates (forget, input and output) to decide what to erase from it, what to write into it and what to pass on as the hidden state at each step.

Last updated: 07 Oct, 2026 · TensorFlow 2 / Keras · NumPy

Vanishing and exploding gradients in RNNs showed a plain RNN losing the early words of a long sentence. The LSTM keeps a second line of memory that information can travel along for many steps without being squeezed through tanh and Wh at every one.

Context switching in long sentences · from Day 8 of the Live NLP series · 12:55 to 15:56

Keeping context across a long sentence

Take "My name is Krish and my friend name is ___". To fill the blank, the network must keep track of who is being talked about: first the speaker, then the friend. The output depends on words far back, and an RNN, with its short-term memory, cannot hold them.

Now take "Krish likes pizza but my friend likes burger". At "friend" the context switches: the sentence is no longer about the speaker. The network should forget the old subject, focus on the new one, and still be able to refer back to the speaker later if the sentence returns to him. Remembering for a long time, and forgetting deliberately when the context changes, is what the name long short-term memory describes.

Reading the LSTM cell

An LSTM cell: the cell state runs along the top from C(t-1) to C(t), first multiplied elementwise by the forget gate f, then added to the product of the input gate i and the candidate values; the concatenated [h(t-1), x(t)] runs along the bottom into the four layers sigma, sigma, tanh and sigma; the new hidden state h(t) is the output gate o times tanh of C(t).

A plain RNN cell has one layer: [ht−1, xt] joined and passed through tanh. The LSTM cell has four interacting layers, three sigmoids and one tanh. The legend of the drawing:

  • A σ box is a layer of neurons with a sigmoid activation; its outputs lie between 0 and 1, so it acts as a gate.
  • A tanh box is a layer with a tanh activation; its outputs lie between −1 and 1.
  • Two lines joining means concatenate; a line splitting means copy.
  • A × or + circle is a pointwise operation: elementwise multiplication or addition of two vectors of the same length.
The memory cell as a conveyor belt · from Day 8 of the Live NLP series · 20:39 to 22:12

Carrying the cell state like a conveyor belt

The line along the top of the cell is the memory cell, also called the cell state Ct. It works like the conveyor belt at an airport: luggage keeps moving along it, and people can take luggage off or put new luggage on. Along the cell state, the LSTM can remove information and add information, and everything else passes through untouched.

The cell state is the network's long-term memory, because it can carry a value across many steps. The hidden state ht along the bottom is the short-term, working state that the gates read and the next layer receives.

The forget gate on two sentences · from Day 8 of the Live NLP series · 26:55 to 30:06

Forgetting with the forget gate

The forget gate decides what to remove from the cell state. The video compares two sentences:

  • "Krish likes pizza but he does not like burger": "he" is still the same person, so there is hardly any context switch. The sigmoid gives values near 1, and multiplying the cell state by values near 1 keeps the information.
  • "Krish likes pizza but his friend likes burger": when "friend" arrives at xt, the previous context must go. The sigmoid gives values near 0, and the pointwise product turns the old values into zeros: [1, 1, 1, 0, 1] times [0, 0, 0, 0, 0] is all zeros.
The forget gate

Wf has two parts, one block of weights for ht−1 and one for xt, and like every LSTM weight it is shared by all the time steps. The gate is learned from data: it is not a similarity score, and it works number by number, so it can clear the dimensions that hold the old subject while keeping others. Keras starts bf at 1 (unit_forget_bias=True), so a new LSTM keeps its memory until training teaches it to forget. The forget gate itself was added to the original 1997 LSTM by Gers, Schmidhuber and Cummins (2000); every LSTM in Keras and PyTorch has it.

Writing new information with the input gate

The second pair of layers decides what to write. A tanh layer proposes candidate values C̃t between −1 and 1, and the input gate it, a sigmoid, chooses how much of each candidate value to let in:

The input gate and the candidate values

The cell state is then updated in one line: forget part of the old state, add the gated candidate. Both products are elementwise (⊙), one number at a time:

The cell state update

For "his friend" the gates need ft near 0 and it near 1 on the dimensions that describe the subject: the old subject is cleared and the new one written. For "he" they need the opposite, ft near 1 and it near 0, so the cell keeps what it had.

Choosing the output with the output gate

The last sigmoid, the output gate, decides which parts of the cell state to expose. The cell state goes through tanh and is multiplied by ot; the result is the new hidden state, which is both the cell's output and its input at the next step:

The output gate and the hidden state

Only ht is filtered. Ct itself goes on to the next step as it is.

Walking through "his friend" with numbers

The cell holds [1, 1, 1, 1, 0] about the first subject, and the candidate for the new subject is [−1, −1, 0, 1, 1]. The code applies both gate settings to the same values.

ExampleRun with NumPy
import numpy as np

C_prev = np.array([1, 1, 1, 1, 0])          # what the cell holds about the first subject
C_tilde = np.array([-1, -1, 0, 1, 1])       # candidate values for a new subject
cases = {"he (same subject)": (np.full(5, 0.98), np.full(5, 0.02)),
         "his friend (new subject)": (np.full(5, 0.02), np.full(5, 0.98))}
for name, (f, i) in cases.items():
    C = f * C_prev + i * C_tilde
    print(f"{name}:")
    print("   f*C(t-1) =", np.round(f * C_prev, 2), " i*C~ =", np.round(i * C_tilde, 2), " C(t) =", np.round(C, 2))
  • "he": f = 0.98 keeps the old values, i = 0.02 lets in almost nothing, and Ct = [0.96, 0.96, 0.98, 1, 0.02] is the old state.
  • "his friend": f = 0.02 clears the old values, i = 0.98 writes the candidate, and Ct = [−0.96, −0.96, 0.02, 1, 0.98] is the new subject.
  • Each gate works per number: the products are elementwise, so the result is a vector of the same length, never a single number.

Running one LSTM step in NumPy

The code runs one full LSTM step on random numbers: 4 inputs, 3 hidden units, the four equations in order. It also checks how Ct changes with Ct−1.

The four layers on [h(t-1), x(t)]

python
v = np.concatenate([h_prev, x_t])         # [h(t-1), x(t)]
f = sigmoid(W_f @ v + b_f)                # forget gate
i = sigmoid(W_i @ v + b_i)                # input gate
C_tilde = np.tanh(W_C @ v + b_C)          # candidate values
o = sigmoid(W_o @ v + b_o)                # output gate

The new cell state and hidden state

python
C = f * C_prev + i * C_tilde              # new cell state, elementwise
h = o * np.tanh(C)                        # new hidden state
ExampleRun with NumPy
import numpy as np

def sigmoid(z):
    return 1 / (1 + np.exp(-z))

rng = np.random.default_rng(42)
D, H = 4, 3
x_t = rng.normal(0, 1, D)                  # the current word's vector
h_prev = rng.normal(0, 0.5, H)             # h(t-1), the short-term state
C_prev = rng.normal(0, 1, H)               # C(t-1), the cell state
W_f, W_i, W_C, W_o = (rng.normal(0, 0.5, (H, H + D)) for _ in range(4))
b_f, b_i, b_C, b_o = np.ones(H), np.zeros(H), np.zeros(H), np.zeros(H)   # forget bias starts at 1

v = np.concatenate([h_prev, x_t])         # [h(t-1), x(t)]
f = sigmoid(W_f @ v + b_f)                # forget gate
i = sigmoid(W_i @ v + b_i)                # input gate
C_tilde = np.tanh(W_C @ v + b_C)          # candidate values
o = sigmoid(W_o @ v + b_o)                # output gate

C = f * C_prev + i * C_tilde              # new cell state, elementwise
h = o * np.tanh(C)                        # new hidden state

for name, val in [("f", f), ("i", i), ("C~", C_tilde), ("o", o), ("C(t-1)", C_prev), ("C(t)", C), ("h(t)", h)]:
    print(f"{name:7s}", np.round(val, 4))

eps = 1e-6                                 # nudge C(t-1) and watch C(t)
dC = [((f * (C_prev + eps * np.eye(H)[k]) + i * C_tilde) - C)[k] / eps for k in range(H)]
print("dC(t)/dC(t-1) on each unit:", np.round(dC, 4), " equals f:", np.allclose(dC, f))
print("100 steps at f = 0.99:", round(0.99 ** 100, 3), "  at f = 0.9:", f"{0.9 ** 100:.2e}", "  0.25 ** 100:", f"{0.25 ** 100:.1e}")
print("LSTM parameters, D=300, H=100:", 4 * (100 * (100 + 300) + 100))
print("Keras LSTM(100) on 40-number embeddings:", 4 * (100 * (100 + 40) + 100))

What the LSTM step shows

  • f, i and o lie between 0 and 1 and C̃ between −1 and 1, as sigmoid and tanh guarantee.
  • C(t) = f ⊙ C(t−1) + i ⊙ C̃: for the first unit, 0.5249 × (−0.3162) + 0.3562 × 0.7613 ≈ 0.105.
  • dC(t)/dC(t−1) equals f on every unit. On the path along the cell state, each step back multiplies the gradient by the forget gate, not by tanh′ and Wh.
  • With f = 0.99 for 100 steps the gradient keeps 0.366 of its size; a plain RNN's sigmoid best case, 0.25¹⁰⁰, is 6.2 × 10⁻⁶¹. At f = 0.9 it falls to 2.66 × 10⁻⁵, so the cell state helps only as long as the gates stay open.
  • An LSTM layer has four times the weights of an RNN layer: 160,400 for D = 300, H = 100, and 56,400 for the 100 units on 40-number embeddings in the fake news model.

LSTM vs RNN

RNNLSTM
Stateshidden state hthidden state ht and cell state Ct
Layers inside the cellone tanhthree sigmoid gates and one tanh
Gradient path back in timethrough tanh′ and Wh at every stepalong the cell state, through ft
Parameters, D = 300, H = 10040,100160,400
Long sentencesearly words fadekept as long as the forget gate stays near 1

Where you use an LSTM

  • Text classification where context spreads over a sentence, such as fake news titles or reviews.
  • Tagging and sequence labelling, often in a bidirectional layer.
  • Time series forecasting with long seasonal patterns.
Watch out. An LSTM reduces vanishing gradients, it does not remove them. If the forget gates sit at 0.9, a hundred steps still shrink the gradient to 2.66 × 10⁻⁵, and very long documents lose their early context. That limit is what attention, and later the transformer, was built to remove.
Try it yourself
  • Set b_f = np.zeros(H) instead of ones and see how much smaller f becomes.
  • In the walk-through, set f = 0.02 and i = 0.02 for "his friend": the cell forgets the old subject but writes almost nothing new.
  • Print 0.95 ** 100 and place it between the f = 0.99 and f = 0.9 results.

This is what real progress feels like.