Speculative Decoding

Turkish equivalent: Spekülatif çözümlemeDomain: Large Language Models

An inference technique in which a cheaper draft model proposes multiple tokens that a larger target model verifies in batches to increase decoding throughput.

Language-Model Context

Speculative decoding lets a cheaper draft model propose several tokens that a target model verifies in a smaller number of target-model steps. When proposal acceptance is high enough, this can improve decoding throughput without changing the target model's output distribution under the algorithm's assumptions.

Performance Boundary

Speedup depends on draft cost, acceptance rate, target-kernel efficiency, sequence shape, and memory behavior. It is not guaranteed for every model pair, and memory-bandwidth limits can remain dominant.

Source

  • https://arxiv.org/abs/2211.17192