Speculative Decoding
An inference technique in which a cheaper draft model proposes multiple tokens that a larger target model verifies in batches to increase decoding throughput.
Language-Model Context
Speculative decoding lets a cheaper draft model propose several tokens that a target model verifies in a smaller number of target-model steps. When proposal acceptance is high enough, this can improve decoding throughput without changing the target model's output distribution under the algorithm's assumptions.
Performance Boundary
Speedup depends on draft cost, acceptance rate, target-kernel efficiency, sequence shape, and memory behavior. It is not guaranteed for every model pair, and memory-bandwidth limits can remain dominant.
Related LLM Concepts
Source
- https://arxiv.org/abs/2211.17192