Architecture matters, but it does not explain everything
An architecture determines how information is represented and processed: which elements can influence one another, in what order computations are performed, and which parts of the context are available when producing an output. In language tasks, these decisions have practical consequences. A translation system, for example, must relate words in the input to words in the output and produce a coherent sequence.
But architecture is not a complete explanation of a system’s capabilities. The training objective, available data, computation, and the way the model is adapted to a task also matter. That is why describing history as a chain in which each design “outperforms” and eliminates the one before it is misleading. What often changes is the balance between benefits and costs, and different techniques can remain useful in different settings.
Machine translation offers a way to examine these changes because it requires solving several problems at once: estimating which expressions in one language correspond to those in another, and producing a target-language sentence. Language modeling provides another case study: learning to assign probabilities to sequences and, in the autoregressive case, predicting which token may come next. These examples make it possible to compare design choices without confusing the architecture with the task.
Before recurrent networks: statistical machine translation
Before the neural translation models discussed in this account, one line of research applied statistical methods to the problem. Peter F. Brown and his colleagues’ paper on statistical methods for machine translation documents an early approach. Rather than treating translation as a fixed list of hand-written linguistic rules, these methods estimated probabilistic relationships from data and sought an appropriate output based on those estimates.
The difference from a neural network is not simply that one approach uses numbers and the other does not: both perform calculations. The relevant difference is how relationships are represented and how the components involved in the task are organized. In a statistical system, the problem is formulated using probability models and procedures for choosing a translation. In neural encoder-decoder systems, by contrast, a network learns internal representations of the input sequence and uses them to generate the output sequence.
This starting point is a reference, not an exhaustive account of all statistical machine translation. A foundational paper can describe a specific contribution, but it is not enough to reconstruct every variant, deployed system, or technical choice that existed in the field. Nor did the move toward neural methods mean that earlier methods disappeared immediately, or that an architecture alone guaranteed better translation on every dataset.
Recurrent encoder-decoder: encoding and generating one step at a time
Sutskever, Vinyals, and Le studied a neural encoder-decoder architecture based on LSTM networks for English-to-French translation. The encoder processes the input sequence, and the decoder produces the output sequence. In the formulation described in their paper, the input is represented by a fixed-length vector that the decoder uses to generate the translation.
This design made it possible to learn a transformation between sequences without manually designing a correspondence for every word. The encoder and decoder could learn useful representations from translation examples. However, the fixed-length vector also concentrated information from the entire input sentence into a representation of limited size. The potential difficulty grows when a sequence is long or contains information that the decoder needs to retrieve precisely: the model must preserve in that representation whatever is needed to produce the output.
Recurrence introduces another trade-off. The state that processes one position depends on the previous state, so a sequence is traversed step by step. This structure offers a natural way to handle variable-length inputs and maintain a state summarizing what has been processed so far. At the same time, it limits how much can be computed in parallel across the sequence, because each step needs the result of the previous one. This is not an abstract drawback: it affects how computation is organized and the cost of training on sequences.
It is important to separate what the paper demonstrated from a broader conclusion. The paper evaluated a particular architecture for English-to-French translation; it does not prove that all recurrent networks compress information in the same way or that a fixed-length vector is always insufficient. The limitation motivates studying mechanisms that can consult the input more directly; it is not proof that recurrence has no value.
How a recurrent encoder-decoder works
A conceptual outline of the translation process described for neural machine translation.
- 01The encoder traverses the elements of the input sentence and updates its recurrent state.
- 02The final fixed-length representation summarizes the input for the decoder.
- 03The decoder produces the translation sequentially, conditioning each step on the context it has received and on what it has generated so far.
Attention: consulting the input without abandoning recurrence
Bahdanau, Cho, and Bengio proposed an attention mechanism for jointly learning to align and translate. Their motivation was to reduce the bottleneck created by representing an entire source sentence with a single fixed-length vector. Rather than requiring the decoder to rely only on that summary, the mechanism lets it search for relevant parts of the input when predicting each target word.
Intuitively, the decoder can assign different importance to different elements of the source sentence at different points in generation. If it needs to produce a particular word, it can rely more heavily on a relevant part of the input than on the rest. This gives the relationship between input and output a more flexible representation and makes explicit an operation related to alignment, which translation must resolve in some way.
The attention in this paper is not yet the Transformer. It is incorporated into an architecture that retains recurrent components: sequence processing and generation are still organized step by step. The important change is that the decoder gets selective access to information in the input instead of receiving it only through a single summary. Attention therefore improves one dimension of the problem—access to the input—without removing the sequential dependence characteristic of recurrent networks.
This distinction helps avoid a common oversimplification: “attention” does not name a single architecture and does not automatically imply full parallelization. It can refer to a mechanism integrated into models that retain recurrence. To understand what changed in a given design, it is necessary to ask where attention is applied, which representations it consults, and which operations still depend on earlier steps.
The Transformer: self-attention and a different computational balance
The Transformer proposes processing sequences with attention mechanisms and no recurrence. Its original paper studied the design on machine-translation tasks and presented an encoder and decoder with self-attention and feed-forward layers. Self-attention allows representations of a sequence to incorporate information from other positions in that same sequence; in the decoder, access is restricted so that predictions cannot use future tokens.
Removing recurrence changes how computation is organized. During training, representations at different positions in a layer can be computed in parallel instead of waiting for the previous recurrent step to finish. This was one of the parallelization advantages attributed to the original design. It does not mean that every operation in the system happens simultaneously: layers are still built in order, and autoregressive output generation still needs to produce one token before it can condition the next on that token.
The change also introduces costs. In self-attention, each position can relate to many other positions. As a result, the work required to compute and store those relationships increases as sequence length grows. The architecture does not eliminate computational cost; it redistributes where that cost is incurred. The original paper compares cost, the number of sequential operations, and the complexity of paths between positions, as well as measuring results on translation tasks. Those comparisons support advantages under the conditions evaluated, not a universal rule for every task or sequence length.
It is also important not to turn the history into an instantaneous transition. The Transformer demonstrated an alternative that could take advantage of parallelization during training, but recurrent networks and other techniques did not become impossible as a result. Architectural choice depends on the task, context length, resources, objective, and training system. The history describes a major change in the balance between context access and sequential execution, not the automatic disappearance of every previous option.
What changes with self-attention
A schematic comparison of architectural properties; it is not a universal performance ranking.
| Aspect | RNN or recurrent encoder-decoder | Transformer |
|---|---|---|
| Dependence between positions | The state at one position depends on the state at the previous position. | Self-attention relates positions within a layer without traversing them through a recurrent state. |
| Training on a sequence | The recurrent traversal limits parallel computation across steps. | Positions within a layer can be processed in parallel during training. |
| Autoregressive generation | The output is generated step by step. | The output is also generated step by step when each token depends on earlier tokens. |
| Access to the input | It may depend on recurrent states or, with attention, consult parts of the source. | Self-attention can combine information from multiple positions; its cost grows with the relationships among them. |
GPT-style autoregressive models: objective and architecture
The GPT line shows why it is necessary to distinguish architecture from training objective. OpenAI’s paper on generative pretraining describes a two-stage process: first, generative pretraining on unlabeled text; then, discriminative fine-tuning for language-understanding tasks. With the autoregressive generative objective, the model learns to predict the next token conditioned on the preceding context. Generation applies that same dependency successively: each new token is added to the context to produce the next one.
In the GPT-style model family, the design centers on the decoder component and autoregressive prediction, rather than necessarily building a separate encoder and decoder for translation. This means that the same modeling interface—continuing a sequence—can serve as a starting point for language tasks. But it does not follow that the architecture alone produces general capabilities. The result also depends on the pretraining text, the scale and procedure of training, and how the downstream task is formulated and evaluated.
Pretraining can provide parameters that are later adapted to a task, but it should not be confused with fine-tuning itself or with evaluation. “It is trained to predict the next token” describes an objective; it is not equivalent to claiming that the model understands every instruction, has up-to-date information, or reliably solves every task. Claims of that kind require results for the specific model and task.
Thus, Transformer and GPT are not synonyms. Transformer names an attention-based architecture; GPT refers to a family of autoregressive models and to an approach involving generative pretraining followed, in the cited foundational work, by adaptation to tasks. An architecture can be used with different objectives, and a training objective does not by itself determine all later uses or behaviors.
From pretraining to a task
A summary of the procedure described in the original paper on generative pretraining.
- 01Pretrain the model on unlabeled text using a generative objective.
- 02Prepare the model for a downstream task through discriminative fine-tuning, as specified in the paper.
- 03Evaluate the system on that task; do not infer its results solely from the architecture’s name or its pretraining objective.
A cross-cutting comparison: what changed and what persisted
The most useful comparison is not a list of winners, but a set of questions. Does the model process a sequence step by step, or can it compute representations for multiple positions in parallel? How does it retrieve information from the input when generating an output? Which objective is optimized? What cost arises as sequence length increases? Which part of the result is due to the design, and which to data, computation, or adaptation? These questions make it possible to analyze current systems without assuming that one isolated innovation explains all their performance.
In RNNs, the recurrent state organizes information and creates a dependency between steps. With attention in a recurrent model, the decoder can consult parts of the input, but recurrence remains. The Transformer removes recurrence and enables more parallelism during training, at the cost of computing attention relationships whose cost grows with the number of positions. In an autoregressive model, even one built with a Transformer, generating a sequence remains step by step: the architecture parallelizes part of training, not the logical dependence of each generated token on those before it.
The need to choose an appropriate objective also persists. An encoder-decoder translation architecture and a next-token prediction model are not necessarily trained to solve the same task in the same way. Statistical and neural translation also differ in more than speed or quality: they organize the estimates and representations that lead to an output differently. Performance comparisons must account for the dataset, protocol, and metric used; it is not rigorous to extrapolate a single measurement to every application.
The following table summarizes general trade-offs. It does not replace reviewing a paper or implementation: different variants within a family can modify these trade-offs, and evaluation conditions can change the result.
A problem-oriented comparison guide
Practical questions to guide a comparison; this does not prescribe a universal architecture.
| If the priority is… | Examine… | A limitation not to overlook |
|---|---|---|
| Understanding information flow in translation | Whether there is an encoder-decoder, how the input is represented, and whether the decoder has attention. | A summary vector or an attention mechanism alone does not guarantee a correct translation. |
| Parallelizing computation during training | Whether layer-wise processing avoids recurrent dependencies between positions. | Autoregressive generation retains sequential dependencies, and attention has costs tied to sequence length. |
| Using pretraining across multiple tasks | Which objective was used, on what data, and what adaptation procedure was applied. | A next-token objective is not enough to attribute general capabilities or reliability. |
| Comparing published results | Which task, data, metric, and conditions were evaluated in the original paper. | A result on translation or language understanding does not automatically generalize to other tasks. |
Conclusion: architecture is one part of the explanation
The path from statistical translation methods to GPT-style autoregressive models is not an inevitable march toward a single way of processing language. The recurrent encoder-decoder learned a transformation between sequences, but summarized the input in a fixed-length vector. Attention introduced selective access to parts of that input while retaining recurrence. The Transformer removed recurrent dependencies from the processing of positions and favored parallelization during training, but retained attention costs and token-by-token generation. GPT combined an autoregressive architecture with generative pretraining and later adaptation, without the objective or architecture alone explaining all observed capabilities.
This history helps frame concrete questions about a current model: what context can it consult, which parts of the computation can be parallelized, how is the output produced, and what objective did it learn? It does not automatically answer whether the model is accurate, safe, or suitable for a particular use. That requires evidence about the system and the task. Separating architecture, training, and evaluation avoids both the simplistic story of successive replacements and the attribution of capabilities to a single innovation.
To continue exploring, connect this account with a fundamentals guide to language models, a comparison of approaches, and an overview of AI techniques. In each case, the key question is the same: what evidence describes the system’s behavior, and what part is an interpretation of its design?
Open questions
- The cited papers are relevant foundational case studies, not an exhaustive history of every statistical, recurrent, or neural variant developed.
- The Transformer’s parallelization and performance advantages are based on the tasks and conditions studied in the original paper; they should not automatically be generalized to every sequence or implementation.
- The account summarizes the objective and procedure in the original GPT paper; by itself, it does not allow conclusions about the behavior of later models or their reliability in specific uses.
- The qualitative comparisons among architectures describe general trade-offs. The best choice depends on the task, data, resources, and implementation.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction