14 papers · ranked by Valyu relevance
Anas Al-okaily, Abdelghani Tbakhi
Data compression is a challenging and increasingly important problem. As the amount of data generated daily continues to increase, efficient transmission and storage has never been more critical. In this study, a novel encoding algorithm is proposed, motivated by the compression of DNA data and associated…
Senik Matinyan, Jan Pieter Abrahams
High-throughput data collection in crystallography poses significant challenges in handling massive amounts of data. Here, we present TERSE, a novel lossless compression algorithm specifically designed for diffraction data. We compare TERSE with the established lossless compression algorithms implemented in gzip, CBF…
Łukasz Roguski, Idoia Ochoa, Mikel Hernaez, Sebastian Deorowicz
The affordability of DNA sequencing has led to the generation of unprecedented volumes of raw sequencing data. These data must be stored, processed, and transmitted, which poses significant challenges. To facilitate this effort, we introduce FaStore, a specialized compressor for FASTQ files. The proposed algorithm does…
Herbert J. Bernstein, Alexei S. Soares, Kimberly Horvat, Jean Jakoncic
New higher-count-rate, integrating, large area X-ray detectors with framing rates as high as 17,400 images per second are beginning to be available. These will soon be used for specialized MX experiments but will require optimal lossy compression algorithms to enable systems to keep up with data throughput. Some…
Sagnik Banerjee, Carson Andorf
Advancement in technology has enabled sequencing machines to produce vast amounts of genetic data, causing an increase in storage demands. Most genomic software utilizes read alignments for several purposes including transcriptome assembly and gene count estimation. Herein we present, ABRIDGE, a state-of-the-art…
Alessio P. Buccino, Olivier Winter, David Bryant, David Feng + 2 more
With the rapid adoption of high-density electrode arrays for recording neural activity, electrophysiology data volumes within labs and across the field are growing at unprecedented rates. For example, a one-hour recording with a 384-channel Neuropixels probe generates over 80 GB of raw data. These large data volumes…
Kavindu Jayasooriya, Sasha P. Jenner, Pasindu Marasinghe, Udith Senanayake + 5 more
Nanopore sequencing is an increasingly central tool for genomics. Despite rapid advances in the field, large data volumes and computational bottlenecks continue to pose major challenges. Here we introduce ex-zd, a new data compression strategy that helps address the large size of raw signal data generated during…
Bálint Balázs, Joran Deschamps, Marvin Albert, Jonas Ries + 1 more
Fluorescence imaging techniques such as single molecule localization microscopy, high-content screening and light-sheet microscopy are producing ever-larger datasets, which poses increasing challenges in data handling and data sharing. Here, we introduce a real-time compression library that allows for very fast (beyond…
Bin Duan, Logan A Walker, Bin Xie, Wei Jie Lee + 3 more
Recent advances in microscopy have pushed imaging data generation to an unprecedented scale. While scientists benefit from higher spatiotemporal resolutions and larger imaging volumes, the increasing data size presents significant storage, visualization, sharing, and analysis challenges. Lossless compression typically…
Luke Staniscia, Yun William Yu
Because of the rapid generation of data, the study of compression algorithms to reduce storage and transmission costs is important to bioinformaticians. Much of the focus has been on sequence data, including both genomes and protein amino acid sequences stored in FASTA files. Current standard practice is to use an…
J.M. Lázaro-Guevara, K.M. Garrido
Undeveloped countries like Guatemala, where access to high-speed internet connections is limited, downloading and sharing Biological information of thousands of Mega Bits is a huge problem for the beginning and development of Bioinformatics. Based on that information is an urgent necessity to find a better way to share…
Sebastian Deorowicz, Joanna Walczyszyn, Agnieszka Debudaj-Grabysz
Bioinformatics databases grow rapidly and achieve values hardly to imagine a decade ago. Among numerous bioinformatics processes generating hundreds of GB is multiple sequence alignments of protein families. Its largest database, i.e., Pfam, consumes 40–230 GB, depending of the variant. Storage and transfer of such…
Shubham Chandak, Kedar Tatwawadi, Srivatsan Sridhar, Tsachy Weissman
Nanopore sequencing provides a real-time and portable solution to genomic sequencing, with long reads enabling better assembly and structural variant discovery than second generation technologies. The nanopore sequencing process generates huge amounts of data in the form of raw current data, which must be compressed to…
Idoia Ochoa, Mikel Hernaez, Rachel Goldfeder, Tsachy Weissman + 1 more
Recent advancements in sequencing technology have led to a drastic reduction in the cost of genome sequencing. This development has generated an unprecedented amount of genomic data that must be stored, processed, and communicated. To facilitate this effort, compression of genomic files has been proposed. Specifically…