17 papers · ranked by Valyu relevance
Guoxuan Xia, Luka Ribar, Paul Balanca
Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft tokens which are then verified in parallel by the LLM. Standard speculative decoding is lossless: its rejection and resampling steps exactly preserve the LLM's sampling distribution. Recent work argues that…
Kimonas Provatas, Aris Karatzikos, Charalampos Koilakos, Michail Patsakis + 5 more
Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose…
Frédéric Berdoz, Peer Rheinboldt, Roger Wattenhofer
Speculative decoding accelerates language model inference by separating generation into fast drafting and parallel verification. Its main limitation is drafter–verifier misalignment, which limits token acceptance and reduces overall effectiveness. While small drafting heads trained from scratch compensate with speed…
Kimonas Provatas, Aris Karatzikos, Charalampos Koilakos, Michail Patsakis + 5 more
Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose…
Zhiyuan Ning, Jiawei Shao, Ruge Xu, Xinfei Guo + 3 more
Speculative decoding has become a widely adopted as an effective technique for lossless inference acceleration when deploying large language models (LLMs). While on-the-fly self-speculative methods offer seamless integration and broad utility, they often fall short of the speed gains achieved by methods relying on…
Clara Mohri, Haim Kaplan, Tal Schuster, Yishay Mansour + 1 more
Transformer language models generate text autoregressively, making inference latency proportional to the number of tokens generated. Speculative decoding reduces this latency without sacrificing output quality, by leveraging a small draft model to propose tokens that the larger target model verifies in parallel. In…
Shuoyang Sun, Chang Da, Hao Fang, Kuofeng Gao + 5 more
Speculative decoding has become a widely adopted technique for accelerating large language model (LLM) inference by drafting multiple candidate tokens and verifying them with a target model in parallel. Its efficiency, however, critically depends on the average accepted length $τ$, i.e., how many draft tokens survive…
Ryan Seah, Marvin Rübenacke, Warren J. Gross
Polar codes achieve channel capacity as block length increases, but this asymptotic advantage comes at the cost of decoding speed: the conventional successive cancellation (SC) algorithm is inherently sequential, which limits its practicality for high-throughput applications. This work addresses that limitation by…
Hao Ding, Nannan Wu, Tianyi Qiu
DNA foundation models such as Evo2 7B adopt hybrid Hyena/attention architectures (Striped-Hyena2) whose single-stream autoregressive decoding is bounded by weight bandwidth at ∼45 tok/s. Speculative decoding on such hybrids faces a systems problem that prior SSM work solves only partially: after a draft is verified…
Maryam Mostafalu, Tommy Clausner, Maxime Ferez, Danila Shelepenkov + 5 more
Attention is a fundamental mechanism enabling the brain to overcome its limited capacity for parallel processing. In non-human primates, invasive electrophysiology has shown that attentional selection operates rhythmically, primarily within the alpha (∼8–12 Hz) and theta (∼4–5 Hz) bands. Whether such finely resolved…
Chu-Jung Wu, Chien-Ying Lin, Yu-Chih Huang, Jun Chen
In modern communication systems, packets with different blocklengths often coexist, presenting new challenges for interference management and decoding. In scenarios where short-packet transmissions must meet strict latency and reliability requirements, conventional interference cancellation decoding strategies may be…
Alexis D MacIntyre, Clément Gaultier, Tobias Goehring
During speech perception, properties of the acoustic stimulus can be reconstructed from the listener’s brain using methods such as electroencephalography (EEG). Most studies employ the amplitude envelope as a target for decoding; however, speech acoustics can be characterised on multiple dimensions, including as…
Pengxi Fu, Zhen Wang, Jianxin Guo, Yushuai Zhang + 4 more
Modern communication systems increasingly leverage multiple information streams-including channel observations, statistical models, and contextual knowledge-to enhance decoding reliability. However, the varying and often unpredictable quality of these sources poses a critical challenge: rigid combination rules fail…
Ibrahim Nawaz, Parv Agarwal, Thomas Heinis
DNA storage is a developing field that uses DNA to archive digital data owing to its superior information density and stability. Although DNA storage has been performed on a significant scale, challenges arise from the synthesis and sequencing of data-encoded oligonucleotides. Synthesis of DNA introduces significant…
Lorenzo Posani
Neural decoding is a powerful approach for inferring which variables are represented in the activity of a population of neurons, with broad applications ranging from basic neuroscience to clinical settings such as brain-computer interfaces. More recently, decoding has also been used as a cross-validated tool for…
Yuhang Wang, Weihua Chen, Linjing Song, Zhiping Xu + 6 more
With the rapid growth of data volume in sensor networks, lossy source coding systems achieve high-efficiency data compression with low distortion under limited transmission bandwidth. However, conventional compression algorithms rely on a two-stage framework with high computational complexity and frequently struggle to…
Guoming Song, Dongming Pi, Shancheng Zhao, Chi Wan Sung
Product codes (PCs) are widely used in high-speed communication systems due to their attractive trade-off between error-correction performance and complexity. To further meet the rapidly growing demand for higher data rates, soft-aided hard-decision decoders (SA-HDDs) have been developed. In this paper, we present a…