Search · four archives
Search · four archives
25 papers · ranked by Valyu relevance
Sebastian Wandelt, Ulf Leser, M. Sohel Rahman
The success of high-throughput sequencing has lead to an increasing number of projects which sequence large populations of a species. Storage and analysis of sequence data is a key challenge in these projects, because of the sheer size of the datasets. Compression is one simple technology to deal with this challenge.…
Guillermo Dufort y Álvarez, Gadiel Seroussi, Pablo Smircich, José Sotelo-Silveira + 2 more
Nanopore sequencing technologies are rapidly gaining popularity, in part, due to the massive amounts of genomic data they produce in short periods of time (up to 8.5 TB of data in less than 72 hours). In order to reduce the costs of transmission and storage, efficient compression methods for this type of data are…
Hirak Sarkar, Rob Patro
The past decade has seen an exponential increase in biological sequencing capacity, and there has been a simultaneous effort to help organize and archive some of the vast quantities of sequencing data that are being generated. While these developments are tremendous from the perspective of maximizing the scientific…
Zhi-An Huang, Zhenkun Wen, Qingjin Deng, Ying Chu + 2 more
'Zexuan Zhu'] Background The rapid progress of high-throughput DNA sequencing techniques has dramatically reduced the costs of whole genome sequencing, which leads to revolutionary advances in gene industry. The explosively increasing volume of raw data outpaces the decreasing disk cost and the storage of huge…
Qiuming Luo, Chao Guo, Yi Jun Zhang, Ye Cai + 1 more
Background With the reduction of gene sequencing cost and demand for emerging technologies such as precision medical treatment and deep learning in genome, it is an era of gene data outbreaks today. How to store, transmit and analyze these data has become a hotspot in the current research. Now the compression algorithm…
Yuanjian Liu, Huihao Luo, Zhijun Han, Yao Hu + 5 more
Framework Authors: ['Yuanjian Liu' 'Huihao Luo' 'Zhijun Han' 'Yao Hu' 'Yehui Yang' 'Kyle Chard' 'Sheng Di' 'Ian Foster' 'Jiesheng Wu'] Abstract—Storing and archiving data produced by nextgeneration sequencing (NGS) is a huge burden for research institutions. Reference-based compression algorithms are effective in…
Bobbie Chern, Idoia Ochoa, Alexandros Manolakos, Albert No + 2 more
'Kartik Venkat' 'Tsachy Weissman'] Abstract—DNA sequencing technology has advanced to a point where storage is becoming the central bottleneck in th e acquisition and mining of more data. Large amounts of data are vital for genomics research, and generic compression tools, while viable, cannot offer the same savings as…
Md Ashiqur Rahman, Abdullah Aman Tutul, Sifat Muhammad Abdullah, Md. Shamsuzzoha Bayzid + 1 more
One of the major tasks of bioinformatics is to collect, analyze and interpret large volumes of biomolecular data. The amount of available genomic data is increasing approximately tenfold every year, at a much faster rate than Moore’s Law for computational power . This advancement in sequencing technologies demands more…
Roye Rozov, Ron Shamir, Eran Halperin
Background Data from large Next Generation Sequencing (NGS) experiments present challenges both in terms of costs associated with storage and in time required for file transfer. It is sometimes possible to store only a summary relevant to particular applications, but generally it is desirable to keep all information…
Lampros Gavalakis, Ioannis Kontoyiannis
The problem of lossless data compression with side information available to both the encoder and the decoder is considered. The finite-blocklength fundamental limits of the best achievable performance are defined, in two different versions of the problem: Reference-based compression, when a single side information…
Linqi Wang, Renpeng Ding, Shixu He, Qinyu Wang + 1 more
Metagenomic data compression is very important as metagenomic projects are facing the challenges of larger data volumes per sample and more samples nowadays. The reference-based compression is a promising method to obtain a high compression ratio. However, existing microbial reference genome databases are not suitable…
Kuntai Du, Yihua Cheng, Peder A. Olsen, Shadi A. Noghabi + 2 more
earth observations Authors: ['Kuntai Du' 'Yihua Cheng' 'Peder A. Olsen' 'Shadi A. Noghabi' 'Ranveer Chandra' 'Junchen Jiang'] With the increasing deployment of earth observation satellite constellations, the downlink (satellite-to-ground) capacity often limits the freshness, quality, and coverage of the imagery data…
Kelvin V. Kredens, Juliano V. Martins, Osmar B. Dordal, Mauri Ferrandin + 4 more
The recent decrease in cost and time to sequence and assemble of complete genomes created an increased demand for data storage. As a consequence, several strategies for assembled biological data compression were created. Vertical compression tools implement strategies that take advantage of the high level of similarity…
Pinghao Li, Shuang Wang, Jihoon Kim, Hongkai Xiong + 3 more
'Lucila Ohno-Machado' 'Xiaoqian Jiang' 'Christos A. Ouzounis'] Genome data are becoming increasingly important for modern medicine. As the rate of increase in DNA sequencing outstrips the rate of increase in disk storage capacity, the storage and data transferring of large genome data are becoming important concerns…
Sebastian Wandelt, Ulf Leser
Modern high-throughput sequencing technologies are able to generate DNA sequences at an ever increasing rate. In parallel to the decreasing experimental time and cost necessary to produce DNA sequences, computational requirements for analysis and storage of the sequences are steeply increasing. Compression is a key…
Carl Kingsford, Rob Patro
Storing, transmitting, and archiving the amount of data produced by next generation sequencing is becoming a significant computational burden. For example, large-scale RNA-seq meta-analyses may now routinely process tens of terabytes of sequence. We present here an approach to biological sequence compression that…
Szymon Grabowski, Tomasz M. Kowalski
Genomes within the same species reveal large similarity, exploited by specialized multiple genome compressors. The existing algorithms and tools are however targeted at large, e.g., mammalian, genomes, and their performance on bacteria strains is mediocre. In this work, we propose MBGC, a specialized genome compressor…
Jisung Park, Jeoggyun Kim, Yeseong Kim, Sungjin Lee + 1 more
Data reduction in storage systems is becoming increasingly important as an effective solution to minimize the management cost of a data center. To maximize data-reduction efficiency, existing post-deduplication delta-compression techniques perform delta compression along with traditional data deduplication and lossless…
Anas Al-okaily, Abdelghani Tbakhi
Data compression is a challenging and increasingly important problem. As the amount of data generated daily continues to increase, efficient transmission and storage has never been more critical. In this study, a novel encoding algorithm is proposed, motivated by the compression of DNA data and associated…
Sheng Di, Jinyang Liu, Kai Zhao, Xin Liang + 21 more
'Zhaorui Zhang' 'Milan Shah' 'Yafan Huang' 'Jiajun Huang' 'Xiaodong Yu' 'Congrong Ren' 'Hanqi Guo' 'Grant Wilkins' 'Dingwen Tao' 'Jiannan Tian' 'Sian Jin' 'Zizhe Jian' 'Daoce Wang' 'Md Hasanur Rahman' 'Boyuan Zhang' 'Jon C. Calhoun' 'Guanpeng Li' 'Kazutomo Yoshii' 'Khalid Ayed Alharthi' 'Franck Cappello'] SHENG DI…
Anders Andreasen, Maria Bonto, Fernando Montero
– This paper presents a framework for optimisation and techno-economic analysis of various pressurisation pathways for CO2 pipeline transportation. The pressurisation pathways include a conventional compression only case from initial to final pressure, a sub-critical compression part followed by cooling, liquefaction…
Yann Collet, Nick Terrell, W. Felix Handte, Danielle Rozenblit + 9 more
In the last few decades, research techniques have improved lossless compression ratios by significantly increasing processing time. However, these techniques have not gained popularity in industry because production systems require high throughput and low resource utilization. Instead, real world improvements in…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
Authors not listed
Cost effective and reliable hydrogen compression remains a challenging barrier in the wide-spread adoption of hydrogen as an energy carrier. The prevailing technology of mechanical compression suffers from several drawbacks, some of which can be addressed by non-mechanical compression strategies (e.g., electrochemical…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…