Natural Language ProcessingNLTK 3.10 · scikit-learn 1.9 · gensim 4.4 · TensorFlow 2 / Keras · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Encoder-decoder (seq2seq) models

An encoder-decoder model (sequence-to-sequence, or seq2seq) is a neural network that reads an input sequence with one recurrent network, the encoder, squeezes it into a context vector, and generates an output sequence of any length from that vector with a second recurrent network, the decoder.

Last updated: 07 Oct, 2026 · TensorFlow 2 / Keras · NumPy

The Types of RNN (one-to-many, many-to-one, many-to-many) lesson drew many-to-many as one output for every input word. Translation does not line up like that: "The cat sat" has three words and its French, "le chat s'est assis", has four. The encoder-decoder reads the whole sentence first and writes the answer afterwards.

The encoder-decoder architecture and its context vector · from the Complete Transformers for NLP One Shot video · 6:53 to 9:54

Reading the input with the encoder

The video revises this architecture before it moves on to transformers. Translating English into French is a sequence-to-sequence task: many words go in, many words come out, and the length of the sentence matters. The encoder is an LSTM. The words x1, x2 and x3 of a sentence go in one per time step, t = 1, 2, 3, after an embedding layer has turned each word into a vector. The outputs the encoder produces along the way are not used.

After the last word, the encoder's final state is the context vector C, which has to stand for the whole sentence. C is passed to the decoder, a second LSTM, which predicts the output words through a softmax layer, and the loss is computed on those predictions. As sentences get longer, one C is no longer enough to hold them, and the BLEU score falls with sentence length. BLEU measures how many word sequences a machine translation shares with human reference translations.

An encoder of three LSTM cells reads The, cat and sat at t = 1, 2 and 3, its per-step outputs unused; its last state is the context vector C. A decoder of five LSTM cells starts from C, reads <start>, le, chat, s'est and assis, and through a softmax writes le, chat, s'est, assis and <end>.
The encoder's states, the context vector and the decoder's next-word distribution

With LSTMs, f and g are LSTM cells and the encoder hands over both its hidden state and its cell state. The model learns the probability of the whole output sentence one word at a time, each word conditioned on C and on the words before it: p(y₁ … yT′ | x₁ … xT) = Π p(yt | C, y₁ … yt−1). This is the formulation of Sutskever, Vinyals and Le (2014), who used deep LSTMs for both halves and found that feeding the source sentence in reverse order improved results. Cho et al. (2014) proposed the same split with a new gated hidden unit, later named the GRU, and called it the RNN encoder-decoder.

Generating the output one word at a time

The decoder needs something to read at its first step, so the target vocabulary gets two extra tokens, <start> (also written <sos> or GO) and <end>. At inference time the loop is:

  1. The encoder reads the whole source sentence and returns C.
  2. The decoder starts from C and reads <start>.
  3. Its softmax gives a probability for every word of the target vocabulary, and the most likely word is written out. This is greedy decoding.
  4. That word becomes the decoder's next input, and the loop repeats until it writes <end> or reaches a length limit.
python
state = encoder(source_tokens)                # C: the encoder's last state
word, output = "<start>", []
for _ in range(max_len):
    probs, state = decoder_step(word, state)  # softmax over the target vocabulary
    word = vocab[probs.argmax()]              # greedy: the most likely word
    if word == "<end>":
        break
    output.append(word)                       # read back in at the next step

Taking only the single best word at each step can lead the decoder into a sentence it cannot finish well. Beam search keeps the few best partial sentences at every step instead; Sutskever et al. decoded with it.

Training with teacher forcing

During training the correct translation is known. Instead of feeding the decoder its own guesses, which are poor early in training, the true previous word is fed at every step. This is teacher forcing. The decoder input is the target shifted one place to the right, with <start> in front, and the label at each step is the target with <end> at the back. Every step has a known answer, so all the steps are scored with cross-entropy in one pass.

ExampleRun on Python 3.12
target = ["le", "chat", "s'est", "assis"]          # "The cat sat" in French

decoder_input = ["<start>"] + target              # the target shifted right by one
decoder_target = target + ["<end>"]               # the word each step must predict

for t, (x, y) in enumerate(zip(decoder_input, decoder_target), start=1):
    print(f"step {t}: decoder reads {x:<8} -> must predict {y}")
  • Five steps for four French words: the extra step is where the decoder learns to stop by predicting <end>.
  • Each input is the label of the step before, so the decoder always sees the correct history while it learns.
  • At inference there is no true history, so the decoder reads its own outputs. The gap between the two is called exposure bias.

Encoding The cat sat in NumPy

The example runs an encoder over the board's embeddings for The, cat and sat, [1, 0, 1, 0], [0, 1, 0, 1] and [1, 1, 1, 1], and then one decoder step. A plain RNN cell, ht = tanh(Wxxt + Whht−1 + b), stands in for the LSTM of the video to keep the arithmetic short. The weights are random with a fixed seed, so the run repeats.

The encoder loop

python
h = np.zeros(3)                          # h0
for word in ["The", "cat", "sat"]:
    h = np.tanh(W_x @ np.array(E[word]) + W_h @ h + b)
context = h                              # C, the last hidden state

The first decoder step

python
s = np.tanh(U_y @ Y[vocab.index("<start>")] + U_s @ context)   # read <start>, start from C
p = np.exp(W_o @ s) / np.exp(W_o @ s).sum()                      # softmax over 6 French words
loss = -np.log(p[vocab.index("le")])                             # teacher forcing: the label is "le"
ExampleRun on NumPy 2.5
import numpy as np

E = {"The": [1, 0, 1, 0], "cat": [0, 1, 0, 1], "sat": [1, 1, 1, 1]}   # the board's embeddings
rng = np.random.default_rng(42)
W_x = rng.normal(0, 0.5, (3, 4))     # input (4 numbers) -> hidden state (3 numbers)
W_h = rng.normal(0, 0.5, (3, 3))     # previous hidden state -> hidden state
b = np.zeros(3)

h = np.zeros(3)                      # h0: the encoder starts empty
for t, word in enumerate(["The", "cat", "sat"], start=1):
    h = np.tanh(W_x @ np.array(E[word]) + W_h @ h + b)
    print(f"t={t} {word:>3}: h = {np.round(h, 4)}")

context = h                          # the last hidden state is the context vector C
print("context vector C =", np.round(context, 4))

# The decoder starts from C and reads <start> as its first input
vocab = ["<start>", "<end>", "le", "chat", "s'est", "assis"]
Y = np.eye(len(vocab))               # one-hot vectors for the French words
U_y = rng.normal(0, 0.5, (3, 6))
U_s = rng.normal(0, 0.5, (3, 3))
W_o = rng.normal(0, 0.5, (6, 3))

s = np.tanh(U_y @ Y[vocab.index("<start>")] + U_s @ context)
logits = W_o @ s
p = np.exp(logits) / np.exp(logits).sum()   # softmax over the 6 French words
print("decoder step 1 probabilities:")
for word, pw in zip(vocab, p):
    print(f"  {word:<8} {pw:.4f}")
print("sum =", round(p.sum(), 4))

loss = -np.log(p[vocab.index("le")])        # teacher forcing: the right word is "le"
print("cross-entropy loss at step 1 =", round(loss, 4))

What the context vector and the probabilities show

  • Three numbers carry the whole sentence: C = [0.0109, −0.9481, 0.2498] is the hidden state after "sat". Only this last state reaches the decoder; the states after "The" and "cat" are not passed on.
  • The six probabilities sum to 1.0. With untrained weights the most likely first word is s'est (0.2717), not le (0.2218).
  • The loss is −ln 0.2218 = 1.5062. Training adjusts every weight of both networks to raise p(le) at this step, and the right label at every other step.

Building the encoder-decoder in Keras

The Keras example for character-level English-to-French translation builds the same two halves from LSTM layers. The encoder keeps only its final states; the decoder starts from them and is trained on teacher-forced inputs.

python
import keras

latent_dim = 256                       # size of the LSTM states h and c
encoder_inputs = keras.Input(shape=(None, num_encoder_tokens))   # one-hot characters
encoder = keras.layers.LSTM(latent_dim, return_state=True)
encoder_outputs, state_h, state_c = encoder(encoder_inputs)
encoder_states = [state_h, state_c]    # keep the states, discard the outputs

decoder_inputs = keras.Input(shape=(None, num_decoder_tokens))   # target, starting with "\t"
decoder_lstm = keras.layers.LSTM(latent_dim, return_sequences=True, return_state=True)
decoder_outputs, _, _ = decoder_lstm(decoder_inputs, initial_state=encoder_states)
decoder_dense = keras.layers.Dense(num_decoder_tokens, activation="softmax")
decoder_outputs = decoder_dense(decoder_outputs)

return_state=True makes the encoder LSTM return its final hidden state h and cell state c next to its output, and the output is discarded. initial_state=encoder_states starts the decoder from those two states. The inputs are one-hot characters: num_encoder_tokens and num_decoder_tokens count the distinct English and French characters.

python
model = keras.Model([encoder_inputs, decoder_inputs], decoder_outputs)
model.compile(optimizer="rmsprop", loss="categorical_crossentropy", metrics=["accuracy"])
model.fit([encoder_input_data, decoder_input_data], decoder_target_data,   # teacher forcing
          batch_size=64, epochs=100, validation_split=0.2)

Every French target starts with a tab character, the start token, and ends with a newline, the end token. decoder_target_data holds the same characters as decoder_input_data moved one step ahead, which is teacher forcing. The full example, Character-level recurrent sequence-to-sequence model, trains on 10,000 English-French sentence pairs and then builds a separate encoder model and decoder model for inference, which feed each predicted character back in.

Encoder-decoder vs a many-to-many RNN

Many-to-many RNN (aligned)Encoder-decoder (seq2seq)
Output lengthsame as the inputany length, ended by
Word orderoutput t belongs to input tfree; the decoder can reorder
What an output step seesthe inputs up to t (both sides with a BiLSTM)the context vector and the words written so far
Typical taskPOS tagging, NERtranslation, summarisation

Where you use encoder-decoder models

  • Machine translation, the task the architecture was designed for.
  • Summarisation and headline writing, where a long input becomes a short output.
  • Speech recognition and chat replies, where audio frames or a message become a sentence of a different length.
Watch out. A model scored with teacher forcing looks better than it is, because it always sees the right previous word. Measure quality with the generation loop, where the decoder reads its own outputs, and include long sentences, where one context vector struggles most.
Try it yourself
  • Add a fourth word to the encoder example (any four numbers) and check that context still has three numbers.
  • Change the label in the loss line from "le" to "s'est" and compare the loss with 1.5062.
  • In the teacher-forcing example, use ["comment", "vas-tu", "?"] for "How are you" and count the decoder steps.

This is what real progress feels like.