Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Recurrent neural network (RNN)

A recurrent neural network (RNN) is a neural network that reads a sequence one element at a time and passes a hidden state from each step to the next, so its output depends on every element it has read so far, in order.

Last updated: 07 Oct, 2026 · TensorFlow 2 / Keras · NumPy

Average Word2Vec turns a whole sentence into one fixed vector, and in doing so it throws away the order of the words. For many tasks the order is the meaning, so the deep learning half of NLP starts with a network that reads words in sequence.

Seeing why word order needs a new network

Text-to-vector methods such as bag of words, TF-IDF and Word2Vec work well for a machine learning model, but the tasks below depend on which word comes after which:

  • Chatbots: a question goes in as a sequence of words and an answer comes out as another.
  • Language translation: Hindi to English, where the output must also be grammatical.
  • Text generation: Gmail's suggestions complete a sentence from the words typed so far.
  • Sentiment analysis: "not good" and "good, not bad" use the same words in a different order.

A plain feed-forward network (an ANN) takes a fixed number of inputs at once and has no notion of first or last. An RNN keeps one small network and feeds it the words one after another, with its own previous output as an extra input. Drawn compactly, that is a neuron with a loop back to itself.

Unrolling an RNN for sentiment analysis · from Day 6 of the Live NLP series · 24:37 to 28:17

Unrolling an RNN over a sentence

The video's example is sentiment analysis: the sentence "The food is good" should give the output Positive. Its words are named x11, x12, x13 and x14: the first index is the sentence (sentence 1) and the second is the word's position.

  • Each word becomes a vector first. x11 goes through Word2Vec and comes out as a vector of d = 300 numbers, the features the network sees.
  • One word per time step. At t = 1 the network takes x11 and produces an output. At t = 2 it takes x12 together with that output, and so on to the last word.
  • The same weights at every step. The weights are initialised once and reused at t = 1, 2, 3 and 4; the forward pass does not change them. A bias is added at every step too.
  • Unrolling is drawing the loop as a chain, one copy of the network per time step. The copies share one set of weights, so the chain is still one network.
An RNN unrolled over the five words The, food, is, very, good: each word vector x11 to x15 enters a tanh cell through the shared weights W, each cell passes its hidden state to the next through the shared weights W_h, h0 is zero, and the last state goes through a sigmoid to give the probability that the review is positive.
Forward propagation in an RNN · from Day 6 of the Live NLP series · 40:22 to 44:24

Writing the forward propagation equations

Sentiment analysis reads many words and gives one answer at the end, so it is a many to one RNN; the outputs of the earlier steps are not used, only passed on. The sentence is now "The food is very good", five words x11 to x15, and its label is Positive.

At t = 1 the input x11 is multiplied by the weights w and passed through the neuron's activation f, giving O1. At t = 2 the network takes x12 times w plus O1 times a second set of weights w′, because the second step also depends on what came before. Every later step follows the same pattern. With vectors these products are matrix multiplications:

The forward pass on the board, one equation per time step

Textbooks (Elman, 1990) write the same recurrence with the hidden state ht for Ot, Wx for w, Wh for w′, a bias b, and tanh as the activation f. The starting state h0 is a vector of zeros, which is why the first equation has no w′ term:

The RNN recurrence (Elman network)

Only the last state feeds the output layer. A yes/no label uses a sigmoid and gives one probability; more than two classes use a softmax. The output layer has its own weights, written wy here so they are not confused with the recurrent ones:

The many-to-one output: P(positive) from the last hidden state

Running the forward pass on "The food is very good"

The code below runs the five steps with NumPy. Random vectors of 4 numbers stand in for the 300-number Word2Vec vectors, and the hidden state has 3 units, so every number fits on a line. The weights are random and untrained, so the probability at the end means nothing yet; training is the subject of Backpropagation through time (BPTT).

The weights shared by every step

python
import numpy as np

rng = np.random.default_rng(42)
words = ["The", "food", "is", "very", "good"]
D, H = 4, 3                        # 4-number vectors stand in for 300-number ones
X = rng.normal(0, 1, (5, D))       # x11 ... x15, one vector per word
W_x = rng.normal(0, 0.5, (H, D))   # the board's w: input weights, shared by every step
W_h = rng.normal(0, 0.5, (H, H))   # the board's w': recurrent weights, shared too
b = np.zeros(H)
w_y = rng.normal(0, 0.5, H)        # output layer on the last state

The loop over the words

One line does each step: the new state is tanh of the current word's input plus the previous state's contribution. The loop runs as many times as the sentence has words.

python
h = np.zeros(H)                    # h0 = 0, so step 1 is tanh(W_x x11 + b)
for t, (word, x) in enumerate(zip(words, X), start=1):
    h = np.tanh(W_x @ x + W_h @ h + b)
    print(f"t={t} {word:5s} h{t} = {np.round(h, 4)}")
ExampleRun with NumPy
import numpy as np

rng = np.random.default_rng(42)
words = ["The", "food", "is", "very", "good"]
D, H = 4, 3                        # 4-number vectors stand in for 300-number ones
X = rng.normal(0, 1, (5, D))       # x11 ... x15, one vector per word
W_x = rng.normal(0, 0.5, (H, D))   # the board's w: input weights, shared by every step
W_h = rng.normal(0, 0.5, (H, H))   # the board's w': recurrent weights, shared too
b = np.zeros(H)
w_y = rng.normal(0, 0.5, H)        # output layer on the last state

h = np.zeros(H)                    # h0 = 0, so step 1 is tanh(W_x x11 + b)
for t, (word, x) in enumerate(zip(words, X), start=1):
    h = np.tanh(W_x @ x + W_h @ h + b)
    print(f"t={t} {word:5s} h{t} = {np.round(h, 4)}")

y_hat = 1 / (1 + np.exp(-(w_y @ h)))
print("P(positive) =", round(float(y_hat), 4))
print("parameters of one RNN layer, D=300, H=100:", 300 * 100 + 100 * 100 + 100)

What the five hidden states show

  • Each ht has 3 numbers between −1 and 1, the range of tanh, whatever the length of the sentence so far.
  • h1 depends on "The" alone, because h0 = 0. Every later state mixes the new word with the state before it, so h5 depends on all five words in order.
  • P(positive) = 0.5946 comes from h5 alone through the sigmoid; with trained weights it would move towards 1 for this sentence.
  • One RNN layer with 300-number inputs and 100 hidden units has 40,100 parameters: 300 × 100 for Wx, 100 × 100 for Wh and 100 biases. The count does not grow with the length of the sentence, because every step reuses the same weights.

RNN vs feed-forward network

Feed-forward network (ANN)RNN
Inputa fixed number of features at oncea sequence, one element per time step
Word orderlost unless built into the featureskept: each step sees the state of the steps before
Weightsone set per layerone set per layer, reused at every time step
Lengthfixedany length; the loop runs once per word
Memorynone between inputsthe hidden state ht

Where you use an RNN

  • Sentiment and text classification where word order matters, read many to one.
  • Tagging each word (parts of speech, named entities), with one output per step.
  • Time series such as next-day sales, where each day's value is one step.
Watch out. A sigmoid output needs a loss that can only be zero or positive. For a yes/no label use binary cross-entropy, −[y log ŷ + (1 − y) log(1 − ŷ)] (Binary cross-entropy). The difference y − ŷ is the error, and it is negative whenever ŷ is above the label.
Try it yourself
  • Change the sentence to four words by using X[:4] and words[:4], and check that the loop runs four times with the same weights.
  • Set W_h to all zeros and run again: each ht then depends on its own word only, which is a feed-forward network applied word by word.
  • Change H = 3 to H = 5 and run again, then work out by hand the parameter count of one RNN layer with D = 300 and H = 50 using the same formula, D × H + H × H + H.

Every expert started right here.