LLM Reference
Concepts & capability filters

direct preference optimization

DPO is a lightweight alignment technique that fine-tunes LLMs directly on pairwise preference data (preferred vs. rejected responses) without a separate reward model or reinforcement learning.

Category
Not classified
Difficulty
Not classified
Aliases
None tracked
Last reviewed
2026-07-02

Key facts

  • It optimizes the policy by maximizing the log-ratio of probabilities between chosen and rejected outputs relative to a reference model.

Models Mentioning direct preference optimization(8)