Concepts & capability filters
Generative Pre-trained Transformer Quantization
GPTQ (Generative Pre-trained Transformer Quantization) is a post-training quantization method that reduces model precision to 4-8 bits while maintaining performance through careful rounding and calibration.
- Category
- Not classified
- Difficulty
- Not classified
- Aliases
- None tracked
- Last reviewed
- 2026-07-02
Key facts
- It enables efficient deployment of large models on consumer hardware with minimal accuracy loss.