Search · four archives
Search · four archives
24 papers · ranked by Valyu relevance
Md Ashiqur Rahman, Abdullah Aman Tutul, Sifat Muhammad Abdullah, Md. Shamsuzzoha Bayzid + 1 more
One of the major tasks of bioinformatics is to collect, analyze and interpret large volumes of biomolecular data. The amount of available genomic data is increasing approximately tenfold every year, at a much faster rate than Moore’s Law for computational power . This advancement in sequencing technologies demands more…
Mathilde Girard, Léa Vandamme, Bastien Cazaux, Antoine Limasset + 1 more
Over the past decade, high-throughput sequencing technologies have transformed sequence bioinformatics into one of the most data-intensive scientific fields. Modern sequencers can now generate tens of terabases of data per day, enabling unprecedented scientific discoveries while simultaneously creating formidable…
Yuanjian Liu, Huihao Luo, Zhijun Han, Yao Hu + 5 more
Framework Authors: ['Yuanjian Liu' 'Huihao Luo' 'Zhijun Han' 'Yao Hu' 'Yehui Yang' 'Kyle Chard' 'Sheng Di' 'Ian Foster' 'Jiesheng Wu'] Abstract—Storing and archiving data produced by nextgeneration sequencing (NGS) is a huge burden for research institutions. Reference-based compression algorithms are effective in…
Linqi Wang, Renpeng Ding, Shixu He, Qinyu Wang + 1 more
Metagenomic data compression is very important as metagenomic projects are facing the challenges of larger data volumes per sample and more samples nowadays. The reference-based compression is a promising method to obtain a high compression ratio. However, existing microbial reference genome databases are not suitable…
Kuntai Du, Yihua Cheng, Peder A. Olsen, Shadi A. Noghabi + 2 more
earth observations Authors: ['Kuntai Du' 'Yihua Cheng' 'Peder A. Olsen' 'Shadi A. Noghabi' 'Ranveer Chandra' 'Junchen Jiang'] With the increasing deployment of earth observation satellite constellations, the downlink (satellite-to-ground) capacity often limits the freshness, quality, and coverage of the imagery data…
Siqi Wu, Yinda Chen, Dong Liu, Zhihai He
Image Compression Authors: ['Siqi Wu' 'Yinda Chen' 'Dong Liu' 'Zhihai He'] In this paper, we study how to synthesize a dynamic reference from an external dictionary to perform conditional coding of the input image in the latent domain and how to learn the conditional latent synthesis and coding modules in an end-to-end…
Jisung Park, Jeoggyun Kim, Yeseong Kim, Sungjin Lee + 1 more
Data reduction in storage systems is becoming increasingly important as an effective solution to minimize the management cost of a data center. To maximize data-reduction efficiency, existing post-deduplication delta-compression techniques perform delta compression along with traditional data deduplication and lossless…
Subhankar Roy, Dilip Kumar Maity, Anirban Mukhopadhyay
Deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) sequence compressors for novel species frequently face challenges when processing wide-scale raw, FASTA, or multi-FASTA structured data. For years, molecular sequence databases have favored the widely used general-purpose Gzip and Zstd compressors. The absence of…
Haichang Yao, Guangyong Hu, Shangdong Liu, Houzhi Fang + 1 more
Since the completion of the Human Genome Project at the turn of the century, there has been an unprecedented proliferation of sequencing data. One of the consequences is that it becomes extremely difficult to store, backup, and migrate enormous amount of genomic datasets, not to mention they continue to expand as the…
Rahul Varki, Christina Boucher
Relative Lempel–Ziv (RLZ) is an effective compression method for large, repetitive collections; however, the fundamental primitives required to elevate it from a passive archival format to a tractable representation for compressed construction have yet to be fully established. In this paper, we introduce an algorithmic…
Hui Sun, Yingfeng Zheng, Haonan Xie, Huidong Ma + 2 more
'Gang Wang'] Background Genomic sequencing reads compressors are essential for balancing high-throughput sequencing short reads generation speed, large-scale genomic data sharing, and infrastructure storage expenditure. However, most existing short reads compressors rarely utilize big-memory systems and duplicative…
Michail Patsakis, Theodore Chronopoulos, Ioannis Mouratidis, Ilias Georgakopoulos-Soares
Genomic data repositories continue to grow as sequencing technologies improve, with the NCBI SRA alone exceeding 47 PB. General-purpose compressors treat bioinformatics files as unstructured byte streams and fail to exploit the structured nature of omics data. We present NYX, a format-aware compression system for…
Tomasz Krokosz, Jarogniew Rykowski, Małgorzata Zajęcka, Robert Brzoza-Woch + 2 more
'Robert Brzoza-Woch' 'Leszek Rutkowski' 'Amitabh Mishra'] Modern, commonly used cryptosystems based on encryption keys require that the length of the stream of encrypted data is approximately the length of the key or longer. In practice, this approach unnecessarily complicates strong encryption of very short messages…
Foad Nazari, Sneh Patel, Melissa LaRocca, Ryan Czarny + 2 more
As sequencing becomes more accessible, there is an acute need for novel compression methods to efficiently store this data. Omics technologies can enhance biomedical research and individualize patient care, but they demand immense storage capabilities, especially when applied to longitudinal studies. Addressing the…
Maria J P Sousa, Armando J Pinho, Diogo Pratas, Inanc Birol
Genomic data, found across a vast array of environments, poses distinct challenges for compression and storage. Effectively compressing genomic sequences necessitates the ability to model their heterogeneous, dynamic, and often incomplete nature, while also accounting for specific genomic features such as a high level…
Luke Staniscia, Yun William Yu
Because of the rapid generation of data, the study of compression algorithms to reduce storage and transmission costs is important to bioinformaticians. Much of the focus has been on sequence data, including both genomes and protein amino acid sequences stored in FASTA files. Current standard practice is to use an…
Sheng Di, Jinyang Liu, Kai Zhao, Xin Liang + 21 more
'Zhaorui Zhang' 'Milan Shah' 'Yafan Huang' 'Jiajun Huang' 'Xiaodong Yu' 'Congrong Ren' 'Hanqi Guo' 'Grant Wilkins' 'Dingwen Tao' 'Jiannan Tian' 'Sian Jin' 'Zizhe Jian' 'Daoce Wang' 'Md Hasanur Rahman' 'Boyuan Zhang' 'Jon C. Calhoun' 'Guanpeng Li' 'Kazutomo Yoshii' 'Khalid Ayed Alharthi' 'Franck Cappello'] SHENG DI…
Christian D. Rask, Daniel E. Lucani
Embedded Applications Authors: ['Christian D. Rask' 'Daniel E. Lucani'] We introduce RAGE, an image compression framework that achieves four generally conflicting objectives: 1) good compression for a wide variety of color images, 2) computationally efficient, fast decompression, 3) fast random access of images with…
Alessio Campanelli, Giulio Ermanno Pibiri, Jason Fan, Rob Patro
We describe lossless compressed data structures for the colored de Bruijn graph (or, c-dBG). Given a collection of reference sequences, a c-dBG can be essentially regarded as a map from k-mers to their color sets. The color set of a k-mer is the set of all identifiers, or colors, of the references that contain the…
Pedro Martin, António Rodrigues, João Ascenso, Maria Paula Queluz
Gaussian Splatting (GS) has emerged as an efficient representation for high-quality 3D reconstruction and novel view synthesis. However, its large model size poses challenges for storage and transmission. While several GS compression solutions have been proposed, their perceptual impact remains poorly understood due to…
Anders Andreasen, Maria Bonto, Fernando Montero
– This paper presents a framework for optimisation and techno-economic analysis of various pressurisation pathways for CO2 pipeline transportation. The pressurisation pathways include a conventional compression only case from initial to final pressure, a sub-critical compression part followed by cooling, liquefaction…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
Authors not listed
Cost effective and reliable hydrogen compression remains a challenging barrier in the wide-spread adoption of hydrogen as an energy carrier. The prevailing technology of mechanical compression suffers from several drawbacks, some of which can be addressed by non-mechanical compression strategies (e.g., electrochemical…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…