13 papers · ranked by Valyu relevance
Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang + 5 more
'Tianyu Liu' 'Wenjie Li' 'Zhifang Sui'] To mitigate the high inference latency stemming from autoregressive decoding in Large Language Models (LLMs), Speculative Decoding has emerged as a novel decoding paradigm for LLM inference. In each decoding step, this method first drafts several future tokens efficiently and…
Minghao Yan, Saurabh Agarwal, Shivaram Venkataraman
Speculative Decoding is a widely used technique to speed up inference for Large Language Models (LLMs) without sacrificing quality. When performing inference, speculative decoding uses a smaller draft model to generate speculative tokens and then uses the target LLM to verify those draft tokens. The speedup provided by…
Hyun Ryu, Eric Kim
Decoding Authors: ['Hyun Ryu' 'Eric Kim'] Inference in Large Language Models (LLMs), such as those used in GPT-3 and LaMDA, has relied heavily on autoregressive decoding, which has yielded effective results. However, with LLMs growing in size and complexity, so has the need for improving inference efficiency. The…
Sanjit Neelam, Daniel Heinlein, Vaclav Cvicek, Akshay Mishra + 1 more
'Reiner Pope'] Speculative decoding (SD) has been shown to reduce the latency of autoregressive decoding (AD) by 2-3× for small batch sizes. However, increasing throughput and therefore reducing the cost per token requires decoding with large batch sizes. Recent work shows that SD can accelerate decoding with large…
Szymon Kobus, Deniz Gündüz
Speculative decoding accelerates large language model inference using a smaller draft model. In this paper, we establish a surprising connection between speculative decoding and channel simulation, which aims at simulating a noisy channel using as few bits as possible. This connection allows us to provide an…
Xiaoxuan Liu, Cade Daniel, Langxiang Hu, Woosuk Kwon + 6 more
Goodput Authors: ['Xiaoxuan Liu' 'Cade Daniel' 'Langxiang Hu' 'Woosuk Kwon' 'Zhuohan Li' 'Xiangxi Mo' 'Alvin Cheung' 'Zhijie Deng' 'Ion Stoica' 'Hao Zhang'] Reducing the inference latency of large language models (LLMs) is crucial, and speculative decoding (SD) stands out as one of the most effective techniques. Rather…
Siru Ouyang, Shuohang Wang, Minhao Jiang, Ming Zhong + 3 more
Distillation Authors: ['Siru Ouyang' 'Shuohang Wang' 'Minhao Jiang' 'Ming Zhong' 'Donghan Yu' 'Jiawei Han' 'Yelong Shen'] Speculative decoding stands as a pivotal technique to expedite inference in autoregressive (large) language models. This method employs a smaller draft model to speculate a block of tokens, which…
Benjamin Spector, Chris Ré
Recent advances with large language models (LLM) illustrate their diverse capabilities. We propose a novel algorithm, staged speculative decoding, to accelerate LLM inference in smallbatch, on-device scenarios. We address the low arithmetic intensity of small-batch inference by improving upon previous work in…
Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Seung‐Yeon Kim + 3 more
'Seung‐Yeon Kim' 'Neha Gupta' 'Aditya Krishna Menon' 'Sanjiv Kumar'] Cascades and speculative decoding are two common approaches to improving language models' inference efficiency. Both approaches involve interleaving models of different sizes, but via fundamentally distinct mechanisms: cascades employ a deferral rule…
Ken R. Duffy, Muriel Médard
In computer communications, discrete data are first channel coded and then modulated into continuous signals for transmission and reception. In a hard detection setting, only demodulated data are provided to the decoder. If soft information on received signal quality is provided, its use can improve decoding accuracy.…
Erdal Arıkan
—Polar coding was conceived originally as a technique for boosting the cutoff rate of sequential decoding, along the lines of earlier schemes of Pinsker and Massey. The key idea i n boosting the cutoff rate is to take a vector channel (either given or artificially built), split it into multiple correlated subchannels…
Mohammed Usman, J. Dunlop
- Luby Transform (LT) codes are a class of fountain codes that have proved to perform very efficiently over the erasure channel. These codes are rateless in the sense that an infinite stream of encoded symbols can be generated on the fly. Furthermore, every encoded symbol is information additive and can contribute in…
Francisco Lázaro Blasco, Gianluigi Liva, Gerhard Bauch
—We present a simple model of inactivation decoding for LT codes which can be used to estimate the decoding complexity as a function of the LT code degree distribution. The model is shown to be accurate in variety of settings of practical importance. The proposed method allows to perform a numerical optimization on the…