8 papers · ranked by Valyu relevance
Dale Zhou, Sharon M. Noh, Nora C. Harhen, Nidhi V. Banavar + 3 more
The ability to discriminate similar visual stimuli has been used as an important index of memory function. This ability is widely thought to be supported by expanding the dimensionality of relevant neural codes, such that neural representations for the similar stimuli are maximally distinct, or “separated.” An…
Alessio P. Buccino, Arjun Sridhar, David Feng, Karel Svoboda + 1 more
The scale of in vivo electrophysiology has expanded in recent years, with simultaneous recordings across thousands of electrodes now becoming routine. These advances have enabled a wide range of discoveries, but they also impose substantial computational demands. Spike sorting, the procedure that extracts spikes from…
Alice Tor, Yuxin Wu, Stephen E Clarke, Lisa Yamada + 2 more
The complexity of neural data changes as the brain processes information during events. Universal lossless compression algorithms, which are broadly applicable and grounded in information theory, identify and exploit redundancies in data in order to compress it to essentially-optimal sizes regardless of underlying…
Marek Kokot, Amitava Roy, Travis J Wheeler, Sebastian Deorowicz
Molecular dynamics (MD) simulations model the physical movements of atoms in biomolecular systems over time, providing atomic-resolution insight into conformational changes, binding events, and dynamic behaviors that cannot be captured by static structures alone. As such, MD simulations are playing an increasingly…
Ibrahim Nawaz, Parv Agarwal, Thomas Heinis
DNA storage is a developing field that uses DNA to archive digital data owing to its superior information density and stability. Although DNA storage has been performed on a significant scale, challenges arise from the synthesis and sequencing of data-encoded oligonucleotides. Synthesis of DNA introduces significant…
Ramy Khabbaz, Jérémy Mateos, Marc Antonini, Serge Kas Hanna
The biochemical processes underlying DNA data storage, including synthesis, amplification, and sequencing, are inherently noisy. Consequently, base-level insertion, deletion, and substitution (IDS) errors, as well as sequence-level dropouts, occur and pose major challenges for reliable data retrieval. Here we introduce…
Adrian Tkachenko, Sepehr Salem, Ayotomiwa Ezekiel Adeniyi, Zülal Bingöl + 6 more
High-throughput sequencing (HTS) enables population-scale genomics but generates massive datasets, creating bottlenecks in storage, transfer, and analysis. FASTQ, the standard format for over two decades, stores one byte per base and one byte per quality score, leading to inefficient I/O, high storage costs, and…
Vojtech Macala, Petr Simecek
Lossless compression and probabilistic sequence modeling are two faces of the same coin: a model that assigns high probability to a sequence can encode it in few bits via arithmetic coding. We exploit this duality to evaluate genomic language models as compressors of DNA, using compression primarily as an objective…