LLM Reference
Concepts & capability filters

transformer architecture

The Transformer architecture consists of encoder and/or decoder stacks with multi-head self-attention, feed-forward layers, and positional encodings for sequence modeling.

Category
Not classified
Difficulty
Not classified
Aliases
None tracked
Last reviewed
2026-07-02

Key facts

  • It enables parallelization and captures context, with decoder-only Transformers powering autoregressive generation in models like GPT and Llama.

Models Mentioning transformer architecture(12)