Encoder-decoder (seq2seq) models
An encoder-decoder model (sequence-to-sequence, or seq2seq) is a neural network that reads an input sequence with one recurrent network, the encoder, squeezes it into a context vector, and generates an output sequence of any length from that vector with a second recurrent network, the decoder.
Last updated: 07 Oct, 2026 · TensorFlow 2 / Keras · NumPy
The Types of RNN (one-to-many, many-to-one, many-to-many) lesson drew many-to-many as one output for every input word. Translation does not line up like that: "The cat sat" has three words and its French, "le chat s'est assis", has four. The encoder-decoder reads the whole sentence first and writes the answer afterwards.
Reading the input with the encoder
The video revises this architecture before it moves on to transformers. Translating English into French is a sequence-to-sequence task: many words go in, many words come out, and the length of the sentence matters. The encoder is an LSTM. The words x1, x2 and x3 of a sentence go in one per time step, t = 1, 2, 3, after an embedding layer has turned each word into a vector. The outputs the encoder produces along the way are not used.
After the last word, the encoder's final state is the context vector C, which has to stand for the whole sentence. C is passed to the decoder, a second LSTM, which predicts the output words through a softmax layer, and the loss is computed on those predictions. As sentences get longer, one C is no longer enough to hold them, and the BLEU score falls with sentence length. BLEU measures how many word sequences a machine translation shares with human reference translations.
With LSTMs, f and g are LSTM cells and the encoder hands over both its hidden state and its cell state. The model learns the probability of the whole output sentence one word at a time, each word conditioned on C and on the words before it: p(y₁ … yT′ | x₁ … xT) = Π p(yt | C, y₁ … yt−1). This is the formulation of Sutskever, Vinyals and Le (2014), who used deep LSTMs for both halves and found that feeding the source sentence in reverse order improved results. Cho et al. (2014) proposed the same split with a new gated hidden unit, later named the GRU, and called it the RNN encoder-decoder.
Generating the output one word at a time
The decoder needs something to read at its first step, so the target vocabulary gets two extra tokens, <start> (also written <sos> or GO) and <end>. At inference time the loop is:
- The encoder reads the whole source sentence and returns C.
- The decoder starts from C and reads <start>.
- Its softmax gives a probability for every word of the target vocabulary, and the most likely word is written out. This is greedy decoding.
- That word becomes the decoder's next input, and the loop repeats until it writes <end> or reaches a length limit.
state = encoder(source_tokens) # C: the encoder's last state
word, output = "<start>", []
for _ in range(max_len):
probs, state = decoder_step(word, state) # softmax over the target vocabulary
word = vocab[probs.argmax()] # greedy: the most likely word
if word == "<end>":
break
output.append(word) # read back in at the next stepTaking only the single best word at each step can lead the decoder into a sentence it cannot finish well. Beam search keeps the few best partial sentences at every step instead; Sutskever et al. decoded with it.
Training with teacher forcing
During training the correct translation is known. Instead of feeding the decoder its own guesses, which are poor early in training, the true previous word is fed at every step. This is teacher forcing. The decoder input is the target shifted one place to the right, with <start> in front, and the label at each step is the target with <end> at the back. Every step has a known answer, so all the steps are scored with cross-entropy in one pass.
target = ["le", "chat", "s'est", "assis"] # "The cat sat" in French
decoder_input = ["<start>"] + target # the target shifted right by one
decoder_target = target + ["<end>"] # the word each step must predict
for t, (x, y) in enumerate(zip(decoder_input, decoder_target), start=1):
print(f"step {t}: decoder reads {x:<8} -> must predict {y}")step 1: decoder reads <start> -> must predict le step 2: decoder reads le -> must predict chat step 3: decoder reads chat -> must predict s'est step 4: decoder reads s'est -> must predict assis step 5: decoder reads assis -> must predict <end>
- Five steps for four French words: the extra step is where the decoder learns to stop by predicting <end>.
- Each input is the label of the step before, so the decoder always sees the correct history while it learns.
- At inference there is no true history, so the decoder reads its own outputs. The gap between the two is called exposure bias.
Encoding The cat sat in NumPy
The example runs an encoder over the board's embeddings for The, cat and sat, [1, 0, 1, 0], [0, 1, 0, 1] and [1, 1, 1, 1], and then one decoder step. A plain RNN cell, ht = tanh(Wxxt + Whht−1 + b), stands in for the LSTM of the video to keep the arithmetic short. The weights are random with a fixed seed, so the run repeats.
The encoder loop
h = np.zeros(3) # h0
for word in ["The", "cat", "sat"]:
h = np.tanh(W_x @ np.array(E[word]) + W_h @ h + b)
context = h # C, the last hidden stateThe first decoder step
s = np.tanh(U_y @ Y[vocab.index("<start>")] + U_s @ context) # read <start>, start from C
p = np.exp(W_o @ s) / np.exp(W_o @ s).sum() # softmax over 6 French words
loss = -np.log(p[vocab.index("le")]) # teacher forcing: the label is "le"import numpy as np
E = {"The": [1, 0, 1, 0], "cat": [0, 1, 0, 1], "sat": [1, 1, 1, 1]} # the board's embeddings
rng = np.random.default_rng(42)
W_x = rng.normal(0, 0.5, (3, 4)) # input (4 numbers) -> hidden state (3 numbers)
W_h = rng.normal(0, 0.5, (3, 3)) # previous hidden state -> hidden state
b = np.zeros(3)
h = np.zeros(3) # h0: the encoder starts empty
for t, word in enumerate(["The", "cat", "sat"], start=1):
h = np.tanh(W_x @ np.array(E[word]) + W_h @ h + b)
print(f"t={t} {word:>3}: h = {np.round(h, 4)}")
context = h # the last hidden state is the context vector C
print("context vector C =", np.round(context, 4))
# The decoder starts from C and reads <start> as its first input
vocab = ["<start>", "<end>", "le", "chat", "s'est", "assis"]
Y = np.eye(len(vocab)) # one-hot vectors for the French words
U_y = rng.normal(0, 0.5, (3, 6))
U_s = rng.normal(0, 0.5, (3, 3))
W_o = rng.normal(0, 0.5, (6, 3))
s = np.tanh(U_y @ Y[vocab.index("<start>")] + U_s @ context)
logits = W_o @ s
p = np.exp(logits) / np.exp(logits).sum() # softmax over the 6 French words
print("decoder step 1 probabilities:")
for word, pw in zip(vocab, p):
print(f" {word:<8} {pw:.4f}")
print("sum =", round(p.sum(), 4))
loss = -np.log(p[vocab.index("le")]) # teacher forcing: the right word is "le"
print("cross-entropy loss at step 1 =", round(loss, 4))t=1 The: h = [ 0.4835 -0.7219 0.4064] t=2 cat: h = [-0.3325 -0.8728 0.154 ] t=3 sat: h = [ 0.0109 -0.9481 0.2498] context vector C = [ 0.0109 -0.9481 0.2498] decoder step 1 probabilities: <start> 0.1381 <end> 0.1605 le 0.2218 chat 0.0848 s'est 0.2717 assis 0.1232 sum = 1.0 cross-entropy loss at step 1 = 1.5062
What the context vector and the probabilities show
- Three numbers carry the whole sentence: C = [0.0109, −0.9481, 0.2498] is the hidden state after "sat". Only this last state reaches the decoder; the states after "The" and "cat" are not passed on.
- The six probabilities sum to 1.0. With untrained weights the most likely first word is s'est (0.2717), not le (0.2218).
- The loss is −ln 0.2218 = 1.5062. Training adjusts every weight of both networks to raise p(le) at this step, and the right label at every other step.
Building the encoder-decoder in Keras
The Keras example for character-level English-to-French translation builds the same two halves from LSTM layers. The encoder keeps only its final states; the decoder starts from them and is trained on teacher-forced inputs.
import keras
latent_dim = 256 # size of the LSTM states h and c
encoder_inputs = keras.Input(shape=(None, num_encoder_tokens)) # one-hot characters
encoder = keras.layers.LSTM(latent_dim, return_state=True)
encoder_outputs, state_h, state_c = encoder(encoder_inputs)
encoder_states = [state_h, state_c] # keep the states, discard the outputs
decoder_inputs = keras.Input(shape=(None, num_decoder_tokens)) # target, starting with "\t"
decoder_lstm = keras.layers.LSTM(latent_dim, return_sequences=True, return_state=True)
decoder_outputs, _, _ = decoder_lstm(decoder_inputs, initial_state=encoder_states)
decoder_dense = keras.layers.Dense(num_decoder_tokens, activation="softmax")
decoder_outputs = decoder_dense(decoder_outputs)return_state=True makes the encoder LSTM return its final hidden state h and cell state c next to its output, and the output is discarded. initial_state=encoder_states starts the decoder from those two states. The inputs are one-hot characters: num_encoder_tokens and num_decoder_tokens count the distinct English and French characters.
model = keras.Model([encoder_inputs, decoder_inputs], decoder_outputs)
model.compile(optimizer="rmsprop", loss="categorical_crossentropy", metrics=["accuracy"])
model.fit([encoder_input_data, decoder_input_data], decoder_target_data, # teacher forcing
batch_size=64, epochs=100, validation_split=0.2)Every French target starts with a tab character, the start token, and ends with a newline, the end token. decoder_target_data holds the same characters as decoder_input_data moved one step ahead, which is teacher forcing. The full example, Character-level recurrent sequence-to-sequence model, trains on 10,000 English-French sentence pairs and then builds a separate encoder model and decoder model for inference, which feed each predicted character back in.
Encoder-decoder vs a many-to-many RNN
| Many-to-many RNN (aligned) | Encoder-decoder (seq2seq) | |
|---|---|---|
| Output length | same as the input | any length, ended by |
| Word order | output t belongs to input t | free; the decoder can reorder |
| What an output step sees | the inputs up to t (both sides with a BiLSTM) | the context vector and the words written so far |
| Typical task | POS tagging, NER | translation, summarisation |
Where you use encoder-decoder models
- Machine translation, the task the architecture was designed for.
- Summarisation and headline writing, where a long input becomes a short output.
- Speech recognition and chat replies, where audio frames or a message become a sentence of a different length.
Related
- Previous: Bidirectional LSTM
- Next: Attention mechanism (Bahdanau and Luong)
- See also: LSTM (long short-term memory)
- Reference: Sutskever, Vinyals and Le, Sequence to Sequence Learning with Neural Networks (2014)
- Add a fourth word to the encoder example (any four numbers) and check that
contextstill has three numbers. - Change the label in the loss line from
"le"to"s'est"and compare the loss with 1.5062. - In the teacher-forcing example, use
["comment", "vas-tu", "?"]for "How are you" and count the decoder steps.
This is what real progress feels like.