Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Heming Xia, Tao Ge, Peiyi Wang, Siqing Chen, Furu Wei, Zhifang Sui · 2023

We propose Speculative Decoding (SpecDec), for the first time ever 1 , to formally study exploiting the idea of speculative execution to accelerate autoregressive (AR) decoding.Speculative Decoding has two innovations: Spec-Drafter -an independent model specially optimized for efficient and accurate drafting -and Spec-Verification -a reliable method for verifying the drafted tokens efficiently in the decoding paradigm.Experimental results on various seq2seq tasks including machine translation and abstractive summarization show our approach can achieve around 5× speedup for the popular Transformer architectures with comparable generation quality to beam search decoding, refreshing the impression that the draft-then-verify paradigm introduces only 1.4×∼2× speedup.In addition to the remarkable speedup, we also demonstrate 3 additional advantages of SpecDec, revealing its practical value for accelerating generative models in real-world applications.Our models and

Read the paper · More papers on PaperTik