22 papers · ranked by Valyu relevance
Byron H. Price, Jeffrey P. Gavornik
While it is universally accepted that the brain makes predictions, there is little agreement about how this is accomplished and under which conditions. Accurate prediction requires neural circuits to learn and store spatiotemporal patterns observed in the natural environment, but it is not obvious how such information…
Yu Yu, Zhihong Sun, Jia Li, Yao Wan + 7 more
Large Language Models (LLMs) are capable of generating syntactically correct and functionally complete programs, greatly streamlining software development. However, recent studies reveal that these programs typically execute substantially slower than human-optimized counterparts. Existing approaches to bridging this…
Veronika Koren, Simone Blanco Malerba, Tilo Schwalger, Stefano Panzeri
The principle of efficient coding posits that sensory cortical networks are designed to encode maximal sensory information with minimal metabolic cost. Despite the major influence of efficient coding in neuroscience, it has remained unclear whether fundamental empirical properties of neural network activity can be…
Yu Yu, Chen Lyu
With the remarkable progress of Code Large Language Models (Code LLMs) in achieving semantic correctness, execution efficiency has become an increasingly important dimension for evaluating their practical utility. However, existing approaches typically treat full programs as a single optimization target during…
Tianhe Wang, Yifan Fang, David Whitney
A paramount challenge for the brain is to precisely model the world and control behavior within the confines of limited encoding capacities. Efficient coding theory posits a unified framework for understanding how neural systems enhance encoding accuracy by tuning to environmental statistics. While this theory has been…
Dong Huang, Jie M. Zhang, Yuhao Qing, Heming Cui
Code generation models have increasingly become integral to aiding software development, offering assistance in tasks such as code completion, debugging, and code translation. Although current research has thoroughly examined the correctness of the code produced by code generation models, a vital aspect — the…
Jesús E. Garca, Verónica A. González-López, Gustavo H. Tasca, Karina Y. Yaginuma + 1 more
In the framework of coding theory, under the assumption of a Markov process $(X_{t})$ on a finite alphabet $A,$ the compressed representation of the data will be composed of a description of the model used to code the data and the encoded data. Given the model, the Huffman’s algorithm is optimal for the number of bits…
Kees Schouhamer Immink, Jos H. Weber, Tuan Thanh Nguyen, Kui Cai + 2 more
The design of low-complexity and efficient constrained codes has been a major research item for many years. This paper reports on a versatile method named concatenated constrained codes for designing efficient fixed-length constrained codes with small complexity. A concatenated constrained code comprises two (or more)…
William Dorrell, Peter E. Latham, Timothy E. J. Behrens, James C. R. Whittington
The efficient coding hypothesis presents a compelling success story for theoretical and systems neuroscience. It marshals a unifying idea, that neural codes can be understood as efficient encodings of natural stimuli, to explain phenomena from across sensory systems, sometimes with exquisite precision. However, similar…
Fajia Sun, Long Qian
DNA has been pursued as a compelling medium for digital data storage during the past decade. While large-scale data storage and random access have been achieved in artificial DNA, the synthesis cost keeps hindering DNA data storage from popularizing into daily life. In this study, we proposed a more efficient paradigm…
Mingyi Huang, Wei Lin, Anna Wang Roe, Yuguo Yu
Understanding how cortical neurons use dynamic firing patterns to represent sensory signals is a central challenge in neuroscience. Decades of research have shown that cortical neuronal activities exhibit high variance, typically quantified by the coefficient of variation (CV), suggesting intrinsic randomness.…
Kun Tu, Dariusz Puchala, Jun Chen, Sadaf Salehkalaibar
In this paper, we address the problem of m-gram entropy variable-to-variable coding, extending the classical Huffman algorithm to the case of coding m-element (i.e., m-grams) sequences of symbols taken from the stream of input data for $m>1$. We propose a procedure to enable the determination of the frequencies of the…
Roland Wittler
To index or compare sequences efficiently, often k-mers, i.e., substrings of fixed length k, are used. For efficient indexing or storage, k-mers are often encoded as integers, e.g., applying some bijective mapping between all possible σ^k^ k-mers and the interval [0, σ^k^ −1], where σ is the alphabet size. In many…
Tomasz Krokosz, Jarogniew Rykowski, Małgorzata Zajęcka, Robert Brzoza-Woch + 2 more
'Robert Brzoza-Woch' 'Leszek Rutkowski' 'Amitabh Mishra'] Modern, commonly used cryptosystems based on encryption keys require that the length of the stream of encrypted data is approximately the length of the key or longer. In practice, this approach unnecessarily complicates strong encryption of very short messages…
Ryosuke Sugiura, Masaaki Nishino, Norihito Yasuda, Yutaka Kamamoto + 1 more
'Takehiro Moriya'] This paper presents an optimal construction of N-bit-delay almost instantaneous fixed-to-variable-length (AIFV) codes, the general form of binary codes we can make when finite bits of decoding delay are allowed. The presented method enables us to optimize lossless codes among a broader class of codes…
Andrew Simmonett, Bernard Brooks, Thomas Darden
Evaluation of noncovalent electrostatic interactions is the dominant bottleneck in classical molecular dynamics simulations, and evaluation of Coulombic matrix elements similarly limits quantum mechanical self consistent field calculations. These difficulties are a result of the Coulomb operator’s slow decay, which…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Canberk İrimağzı, Yusuf Uslan, Ahmed Hareedy
—From the information-theoretic perspective, DNA strands serve as a storage medium for 4-ary data over the alphabet {A, T, G, C}. DNA data storage promises formidable information density, long-term durability, and ease of replicability. However, information in this intriguing storage technology might be corrupted…
Zhimian Hao, Chonghuan Zhang, Alexei Lapkin
We propose a workflow for reduction in the time required for data generation during generation of statistical digital twins. This methodology is particularly relevant for real-world engineering problems when data generation is expensive. A prerequisite for building surrogates is sufficient input/output data, whereas…
Authors not listed
Real-world datasets in chemical engineering and bioengineering processes--such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials--can often be unlabelled or disorganized, rendering the training of existing supervised learning models ineffective at learning the…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Authors not listed
The Hidden Subgroup Problem (HSP) unifies several landmark quantum algorithms, yet systematic exploration of its variants and modern applications has slowed. This paper revives HSP-based algorithm design by examining new group structures with direct relevance to post-quantum cryptography, lattice problems, and…