SpecVocab: Speculative Decoding with a Speculative Vocabulary
Researchers introduce SpecVocab, a method that accelerates language model inference by selecting a dynamic vocabulary subset at each decoding step. This approach achieves higher acceptance lengths and up to an 8.1% increase in average throughput compared to the state-of-the-art EAGLE-3 method, without compromising output quality.
Why it matters: SpecVocab represents a meaningful advance in speculative decoding, offering improved efficiency for large language model inference.
Full story at: arXiv Computation and Language ↗