Concepts & capability filters
direct preference optimization
DPO is a lightweight alignment technique that fine-tunes LLMs directly on pairwise preference data (preferred vs. rejected responses) without a separate reward model or reinforcement learning.
- Category
- Not classified
- Difficulty
- Not classified
- Aliases
- None tracked
- Last reviewed
- 2026-07-02
Key facts
- It optimizes the policy by maximizing the log-ratio of probabilities between chosen and rejected outputs relative to a reference model.