Skip to main content
ZICQ

Wiki AI Concepts

Transformer Architecture

AI Concepts
Aliases: transformer self-attention model ·2026-09-01

Transformer Architecture

The Transformer is a neural network architecture introduced by Google in the 2017 paper Attention Is All You Need. It completely changed NLP and is the foundation of modern large language models.

Core Idea

The core of the Transformer is the self-attention mechanism. Unlike RNNs, the Transformer relies entirely on attention to model dependencies between any two positions in a sequence. This gives two key advantages:

  • Parallel training: Unlike RNNs, the Transformer can process the entire sequence at once.
  • Long-range dependencies: The path length between any two tokens is O(1), avoiding the vanishing-gradient problem of RNNs.

Main Components

A standard Transformer block consists of:

  1. Multi-Head Self-Attention: Lets the model attend to different representational subspaces simultaneously.
  2. Feed-Forward Network: A position-wise non-linear transformation.
  3. Residual Connection + LayerNorm: Stabilizes training of deep networks.

Variants

  • Encoder-only (e.g. BERT): Suited for understanding tasks (classification, extraction).
  • Decoder-only (e.g. GPT): Suited for generation (continuation, chat).
  • Encoder-Decoder (e.g. T5, BART): Suited for translation, summarization, and other seq2seq tasks.

Further Reading