20 papers · ranked by Valyu relevance
James K. Bonfield, Matthew V. Mahoney, Michael Gormley
Storage and transmission of the data produced by modern DNA sequencing instruments has become a major concern, which prompted the Pistoia Alliance to pose the SequenceSqueeze contest for compression of FASTQ files. We present several compression entries from the competition, Fastqz and Samcomp/Fqzcomp, including the…
Anthony J. Cox, Markus Bauer, Tobias Jakobi, Giovanna Rosone
The Burrows-Wheeler transform (BWT) is the foundation of many algorithms for compression and indexing of text data, but the cost of computing the BWT of very large string collections has prevented these techniques from being widely applied to the large sets of sequences often encountered as the outcome of DNA…
W Timothy J White, Michael D Hendy
Background Publicly available DNA sequence databases such as GenBank are large, and are growing at an exponential rate. The sheer volume of data being dealt with presents serious storage and data communications problems. Currently, sequence data is usually kept in large "flat files," which are then compressed using…
Gaëtan Benoit, Claire Lemaitre, Dominique Lavenier, Guillaume Rizk
Motivation: Data volumes generated by next-generation sequencing technologies is now a major concern, both for storage and transmission. This triggered the need for more efficient methods than general purpose compression tools, such as the widely used gzip. Most reference-free tools developed for NGS data compression…
Pamela Vinitha Eric, Gopakumar Gopalakrishnan, Muralikrishnan Karunakaran
'Muralikrishnan Karunakaran'] This paper proposes a seed based lossless compression algorithm to compress a DNA sequence which uses a substitution method that is similar to the LempelZiv compression scheme. The proposed method exploits the repetition structures that are inherent in DNA sequences by creating an offline…
Jeremy John Selva, Xin Chen, James C. Nelson
Next-generation sequencing (NGS) technologies permit the rapid production of vast amounts of data at low cost. Economical data storage and transmission hence becomes an increasingly important challenge for NGS experiments. In this paper, we introduce a new non-reference based read sequence compression tool called…
J.M. Lázaro-Guevara, K.M. Garrido
Undeveloped countries like Guatemala, where access to high-speed internet connections is limited, downloading and sharing Biological information of thousands of Mega Bits is a huge problem for the beginning and development of Bioinformatics. Based on that information is an urgent necessity to find a better way to share…
Anas Al-okaily, Abdelghani Tbakhi
Data compression is a challenging and increasingly important problem. As the amount of data generated daily continues to increase, efficient transmission and storage have never been more critical. In this study, a novel encoding algorithm is proposed, motivated by the compression of DNA data and associated…
Anas Al-okaily, Abdelghani Tbakhi
Data compression is a challenging and increasingly important problem. As the amount of data generated daily continues to increase, efficient transmission and storage has never been more critical. In this study, a novel encoding algorithm is proposed, motivated by the compression of DNA data and associated…
Foad Nazari, Sneh Patel, Melissa LaRocca, Ryan Czarny + 2 more
As sequencing becomes more accessible, there is an acute need for novel compression methods to efficiently store this data. Omics technologies can enhance biomedical research and individualize patient care, but they demand immense storage capabilities, especially when applied to longitudinal studies. Addressing the…
Shakeela Bibi, Javed Iqbal, Adnan Iftekhar, Mir Hassan
Biological data mainly comprises of Deoxyribonucleic acid (DNA) and protein sequences. These are the biomolecules which are present in all cells of human beings. Due to the self-replicating property of DNA, it is a key constitute of genetic material that exist in all breathing creatures. This biomolecule (DNA)…
Neri Merhav
—This article stands as a tribute to the enduring legacy of Jacob Ziv and his landmark contributions to information theory. Specifically, it delves into the groundbreaking individualsequence approach – a cornerstone of Ziv's academic pursuits. Together with Abraham Lempel, Ziv pioneered the renowned Lempel-Ziv (LZ)…
Shengwang Du, Junyi Li, Naizheng Bian, Ruslan Kalendar
The development of high-throughput sequencing technology has generated huge amounts DNA data. Many general compression algorithms are not ideal for compressing DNA data, such as the LZ77 algorithm. On the basis of Nour and Sharawi’s method,we propose a new, lossless and reference-free method to increase the compression…
Fajia Sun, Long Qian
DNA has been pursued as a compelling medium for digital data storage during the past decade. While large-scale data storage and random access have been achieved in artificial DNA, the synthesis cost keeps hindering DNA data storage from popularizing into daily life. In this study, we proposed a more efficient paradigm…
Luke Staniscia, Yun William Yu
Because of the rapid generation of data, the study of compression algorithms to reduce storage and transmission costs is important to bioinformaticians. Much of the focus has been on sequence data, including both genomes and protein amino acid sequences stored in FASTA files. Current standard practice is to use an…
I Made
— People tend to store a lot of files inside theirs storage. When the storage nears it limit, they then try to reduce those files size to minimum by using data compression software. In this paper we propose a new algorithm for data compression, called jbit encoding (JBE). This algorithm will manipulates each bit of…
Mario Mastriani
— A new binary (bit-level) lossless compression catalyst method based on a modular arithmetic, called Binary Allocation via Modular Arithmetic (BAMA), has been introduced in this paper. In other words, BAMA is for storage and transmission of binary sequences, digital signal, images and video, also streaming and all…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
Authors not listed
Sequence is the critical determinant of macromolecular function, yet current polymer design approaches often optimize monomer composition and ratios while ignoring sequence. This creates poorly defined design spaces for active learning that miss the vast combinatorial landscape of sequence possibilities. We introduce…
Authors not listed
This paper presents a simplified model of iterative compound optimization in drug/agrochemical discovery. Compounds are represented as binary strings, with project evolution simulated through random bit changes. The model reproduces key statistical features of real projects, including activity distributions and…