21 Seq2Seq
Context
"Sequence → sequence" tasks (translation) used to need elaborate pipelines with hand-built alignment. Sutskever, Vinyals and Le do it with a single neural network.
The idea and the mechanism
The encoder (an LSTM) reads the input and compresses it into one fixed-length vector — its final hidden state. The decoder (an LSTM) generates the output from that vector word by word: each next word comes from the vector plus whatever has already been generated. It is trained on (input, output) pairs by next-token prediction. The trick: reversing the word order of the source lifted quality noticeably.
probability Autoregressive factorization and the bottleneck
The model factorizes the probability of the whole output sequence into a product of per-step ones (the chain rule for probabilities):
where c is the context vector (the encoder's final state). At each step the decoder takes c and what it has generated so far, and emits a distribution over the next word.
The bottleneck. The entire input x is squeezed into a single fixed-size vector c. The longer the sentence, the more information has to be crammed into the same dimensionality → losses grow with length. Reversing the source helps optimization: the first words of the source and the first words of the translation end up close together in the dependency graph, so the gradient path is shorter — hence the noticeable BLEU gain. But the fundamental bottleneck remains, and attention (#22) will remove it.
PyTorch An LSTM encoder-decoder (sketch)
def seq2seq(src, tgt, enc, dec, embed):
_, (h, c) = enc(embed(src)) # compress the input into a context (h, c)
out, _ = dec(embed(tgt), (h, c)) # the decoder generates from the context
return out # logits over the vocabulary at each step
Why it matters
It cemented the encoder-decoder paradigm on which translation, summarization, dialogue and later multimodal models were built. And its weak spot — the fixed vector — directly produced the next step, attention, and after that the Transformer.
Connections
Both bricks of seq2seq are LSTMs. The ability to hold context over many steps (which the LSTM provided) is a precondition for one vector to represent a whole sentence at all.
A direct continuation: instead of a single vector, the decoder gets access to all the encoder states and chooses for itself what to look at. The bottleneck disappears — and that is the first step towards attention as such.
The Transformer keeps the encoder-decoder idea but throws out recurrence: instead of sequential compression, parallel self-attention. Same paradigm, radically different mechanism.
Questions worth asking
Why does reversing the source help, while reversing the target does not?
Reversing the source shortens the path between the start of the input and the start of the output: the first word of the translation depends on the first words of the original, and now they sit next to each other in the reversed sequence → a shorter gradient path, easier optimization. The target is generated left to right regardless, so reversing it buys nothing. This is a pure optimization trick, not a linguistic one.
If the bottleneck is a known flaw, why study seq2seq at all rather than jump straight to attention?
Because attention is a layer on top of seq2seq, not a replacement: the encoder-decoder skeleton, autoregressive generation and training on pairs all stayed. Understanding the bottleneck is understanding why attention is needed. Without seq2seq, attention looks like an arbitrary trick rather than the fix for a specific problem.
How does the decoder "know" when to stop?
A special end-of-sequence token (<EOS>) is added to the vocabulary; the decoder is trained to emit it when the translation is finished, and generation halts. At inference you normally use beam search — keeping several of the best hypotheses instead of greedily taking the most probable word at each step.
What to read in the original
Read the core part — the encoder-decoder skeleton and the reversal trick; this is the foundation on which attention and the Transformer are built.