Concepts & capability filters
transformer architecture
The Transformer architecture consists of encoder and/or decoder stacks with multi-head self-attention, feed-forward layers, and positional encodings for sequence modeling.
- Category
- Not classified
- Difficulty
- Not classified
- Aliases
- None tracked
- Last reviewed
- 2026-07-02
Key facts
- It enables parallelization and captures context, with decoder-only Transformers powering autoregressive generation in models like GPT and Llama.