LLM Reference
Concepts & capability filters
inference_optimization

Speculative decoding

Speculative decoding speeds up inference by having a small draft model propose several tokens that the target model verifies in parallel, keeping the ones it would have produced anyway.

Category
inference_optimization
Difficulty
Not classified
Aliases
None tracked
Last reviewed
2026-07-02

Key facts

  • Whether output is identical or only near-identical to standard decoding depends on the verifier's acceptance scheme; it is a latency optimization, transparent to how you call the API.

Models Mentioning Speculative decoding(3)