18 papers · ranked by Valyu relevance
Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang + 5 more
'Tianyu Liu' 'Wenjie Li' 'Zhifang Sui'] To mitigate the high inference latency stemming from autoregressive decoding in Large Language Models (LLMs), Speculative Decoding has emerged as a novel decoding paradigm for LLM inference. In each decoding step, this method first drafts several future tokens efficiently and…
Kimonas Provatas, Aris Karatzikos, Charalampos Koilakos, Michail Patsakis + 5 more
Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose…
Kimonas Provatas, Aris Karatzikos, Charalampos Koilakos, Michail Patsakis + 5 more
Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose…
Minghao Yan, Saurabh Agarwal, Shivaram Venkataraman
Speculative Decoding is a widely used technique to speed up inference for Large Language Models (LLMs) without sacrificing quality. When performing inference, speculative decoding uses a smaller draft model to generate speculative tokens and then uses the target LLM to verify those draft tokens. The speedup provided by…
Hyun Ryu, Eric Kim
Decoding Authors: ['Hyun Ryu' 'Eric Kim'] Inference in Large Language Models (LLMs), such as those used in GPT-3 and LaMDA, has relied heavily on autoregressive decoding, which has yielded effective results. However, with LLMs growing in size and complexity, so has the need for improving inference efficiency. The…
Szymon Kobus, Deniz Gündüz
Speculative decoding accelerates large language model inference using a smaller draft model. In this paper, we establish a surprising connection between speculative decoding and channel simulation, which aims at simulating a noisy channel using as few bits as possible. This connection allows us to provide an…
Xiaoxuan Liu, Cade Daniel, Langxiang Hu, Woosuk Kwon + 6 more
Goodput Authors: ['Xiaoxuan Liu' 'Cade Daniel' 'Langxiang Hu' 'Woosuk Kwon' 'Zhuohan Li' 'Xiangxi Mo' 'Alvin Cheung' 'Zhijie Deng' 'Ion Stoica' 'Hao Zhang'] Reducing the inference latency of large language models (LLMs) is crucial, and speculative decoding (SD) stands out as one of the most effective techniques. Rather…
Hao Ding, Nannan Wu, Tianyi Qiu
DNA foundation models such as Evo2 7B adopt hybrid Hyena/attention architectures (Striped-Hyena2) whose single-stream autoregressive decoding is bounded by weight bandwidth at ∼45 tok/s. Speculative decoding on such hybrids faces a systems problem that prior SSM work solves only partially: after a draft is verified…
Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Seung‐Yeon Kim + 3 more
'Seung‐Yeon Kim' 'Neha Gupta' 'Aditya Krishna Menon' 'Sanjiv Kumar'] Cascades and speculative decoding are two common approaches to improving language models' inference efficiency. Both approaches involve interleaving models of different sizes, but via fundamentally distinct mechanisms: cascades employ a deferral rule…
Enzhuo Zhang, Syed Ishtiaque Ahmed
Recent interdisciplinary research has raised interest in whether non-classical physical principles may impose fundamental constraints on how neural information can be observed, extracted, or decoded. While conventional neuroscience models neural signaling primarily through classical electrochemical processes, a growing…
Cameron Higgins, Mats W.J. van Es, Andrew Quinn, Diego Vidaurre + 1 more
Decoding of high temporal resolution, stimulus-evoked neurophysiological data is increasingly used to test theories about how the brain processes information. However, a fundamental relationship between the frequency spectra of the neural signal and the subsequent decoding accuracy timecourse is not widely recognised.…
Ethan M. Meyers
Neural decoding is a powerful method to analyze neural activity. However, the code needed to run a decoding analysis can be complex, which can present a barrier to using the method. In this paper we introduce a package that makes it easy to perform decoding analyses in the R programing language. We describe how the…
Sean M. Perkins, Elom A. Amematsro, John P. Cunningham, Qi Wang + 1 more
Decoders for brain-computer interfaces (BCIs) assume constraints on neural activity, chosen to reflect scientific beliefs while yielding tractable computations. Recent scientific advances suggest that the true constraints on neural activity, especially its geometry, may be quite different from those assumed by most…
Svitlana Matsenko, Oleksiy Borysenko, Sandis Spolitis, Aleksejs Udalcovs + 6 more
'Aleksejs Udalcovs' 'Lilita Gegere' 'Aleksandr Krotov' 'Oskars Ozolins' 'Vjaceslavs Bobrovs' 'Song-Nam Hong' 'T. Aaron Gulliver'] Forward error correction (FEC) codes combined with high-order modulator formats, i.e., coded modulation (CM), are essential in optical communication networks to achieve highly efficient and…
Mingkang Li, Ruixue Wang, Guihua Wan, Yuqi Yang + 1 more
Calcium imaging has gained extensive application in neural decoding tasks because of its high precision in observing cortical neural activity. Nevertheless, the immense data volume and complexity of automated signal extraction algorithms in calcium imaging result in significant delays in extracting neuronal calcium…
Niklas Gassner, Julia Lieb, Abhinaba Mazumder, Michael Schaller
In this paper, we present a framework for generic decoding of convolutional codes, which allows us to do cryptanalysis of code-based systems that use convolutional codes as public keys. We then apply this framework to information set decoding, study success probabilities and give tools to choose variables. Finally, we…
Alireza Tasdighi, Mansoor Yousefi, Jun Chen
Weighted belief propagation (WBP) for the decoding of linear block codes is considered. In WBP, the Tanner graph of the code is unrolled with respect to the iterations of the belief propagation decoder. Then, weights are assigned to the edges of the resulting recurrent network and optimized offline using a training…
Hanqi Tang, Ruobin Zheng, Zongpeng Li, Keping Long + 3 more
'Shenghao Yang' 'Kenneth Shum'] In complex network environments, there always exist heterogeneous devices with different computational powers. In this work, we propose a novel scalable random linear network coding (RLNC) framework based on embedded fields, so as to endow heterogeneous receivers with different decoding…