12 papers · ranked by Valyu relevance
Fabio Cunial, Olgert Denas, Djamal Belazzougui
Fast, lightweight methods for comparing the sequence of ever larger assembled genomes from ever growing databases are increasingly needed in the era of accurate long reads and pan-genome initiatives. Matching statistics is a popular method for computing whole-genome phylogenies and for detecting structural…
Dale Zhou, Sharon M. Noh, Nora C. Harhen, Nidhi V. Banavar + 3 more
The ability to discriminate similar visual stimuli has been used as an important index of memory function. This ability is widely thought to be supported by expanding the dimensionality of relevant neural codes, such that neural representations for the similar stimuli are maximally distinct, or “separated.” An…
Kavindu Jayasooriya, Sasha P. Jenner, Pasindu Marasinghe, Udith Senanayake + 5 more
Nanopore sequencing is an increasingly central tool for genomics. Despite rapid advances in the field, large data volumes and computational bottlenecks continue to pose major challenges. Here we introduce ex-zd, a new data compression strategy that helps address the large size of raw signal data generated during…
Bin Duan, Logan A Walker, Bin Xie, Wei Jie Lee + 3 more
Recent advances in microscopy have pushed imaging data generation to an unprecedented scale. While scientists benefit from higher spatiotemporal resolutions and larger imaging volumes, the increasing data size presents significant storage, visualization, sharing, and analysis challenges. Lossless compression typically…
Alice Tor, Yuxin Wu, Stephen E Clarke, Lisa Yamada + 2 more
The complexity of neural data changes as the brain processes information during events. Universal lossless compression algorithms, which are broadly applicable and grounded in information theory, identify and exploit redundancies in data in order to compress it to essentially-optimal sizes regardless of underlying…
Nathaniel Imel, Jennifer Culbertson, Simon Kirby, Noga Zaslavsky
It has recently been theorized that languages evolve under pressure to attain near-optimal lossy compression of meanings into words. While this theory has been supported by broad crosslinguistic empirical evidence, it remains largely unknown what cognitive mechanisms may drive the cultural evolution of language toward…
Sebastian Deorowicz, Adam Gudyś
The introduction of Deep Minds’ Alpha Fold 2 enabled prediction of protein structures at unprecedented scale. AlphaFold Protein Structure Database and ESM Metagenomic Atlas contain hundreds of millions of structures stored in CIF and/or PDB formats. When compressed with a general-purpose utility like gzip, this…
Herbert J. Bernstein, Alexei S. Soares, Kimberly Horvat, Jean Jakoncic
New higher-count-rate, integrating, large area X-ray detectors with framing rates as high as 17,400 images per second are beginning to be available. These will soon be used for specialized MX experiments but will require optimal lossy compression algorithms to enable systems to keep up with data throughput. Some…
Junjie Tong, Miaoshan Lu, Bichen Peng, Shaowei An + 2 more
The size of high-resolution mass spectrometry (HRMS) data has been increasing significantly. Several lossy compressors have been developed for higher compression rate. Currently, a comprehensive evaluation of what and how MS data (m/z and intensities) with precision losses would affect data processing is absent.…
Fajia Sun, Long Qian
DNA has been pursued as a compelling medium for digital data storage during the past decade. While large-scale data storage and random access have been achieved in artificial DNA, the synthesis cost keeps hindering DNA data storage from popularizing into daily life. In this study, we proposed a more efficient paradigm…
Ibrahim Nawaz, Parv Agarwal, Thomas Heinis
DNA storage is a developing field that uses DNA to archive digital data owing to its superior information density and stability. Although DNA storage has been performed on a significant scale, challenges arise from the synthesis and sequencing of data-encoded oligonucleotides. Synthesis of DNA introduces significant…
Andreas L. Gimpel, Alex Remschak, Wendelin J. Stark, Reinhard Heckel + 1 more
A wide range of codecs with vastly different error-correction approaches have been proposed and implemented for DNA data storage to date. However, while many codecs claim to provide superior performance, no studies have systematically benchmarked codec implementations to establish the current state-of-the-art in DNA…