17 papers · ranked by Valyu relevance
Chowdhury Mofizur Rahman, Mahbub E. Sobhani, Anika Tasnim Rodela, Swakkhar Shatabda
'Swakkhar Shatabda'] Abstract—Text compression shrinks textual data while keeping crucial information, eradicating constraints on storage, bandwidth, and computational efficacy. The integration of lossless compression techniques with transformer-based text decompression has received negligible attention, despite the…
A. E. Allahverdyan, Andranik Khachatryan
A text written using symbols from a given alphabet can be compressed using the Huffman code, which minimizes the length of the encoded text. It is necessary, however, to employ a text-specific codebook, i.e. the symbol-codeword dictionary, to decode the original text. Thus, the compression performance should be…
Beniamin Stecuła, Kinga Stecuła, Adrian Kapczyński, Antonio Puliafito
'Antonio Puliafito'] The goal of the research was to study the possibility of using the planned language Esperanto for text compression, and to compare the results of the text compression in Esperanto with the compression in natural languages, represented by Polish and English. The authors performed text compression in…
Sören Dréano, Derek Molloy, Noel Murphy
This work introduces Llamazip, a novel lossless text compression algorithm based on the predictive capabilities of the LLaMA3 language model. Llamazip achieves significant data reduction by only storing tokens that the model fails to predict, optimizing storage efficiency without compromising data integrity. Key…
Mohammad Hosseini
—Today, with the growing demands of information storage and data transfer, data compression is becoming increasingly important. Data Compression is a technique which is used to decrease the size of data. This is very useful when some huge files have to be transferred over networks or being stored on a data storage…
Swathi Shree Narashiman, Nitin Chandrachoodan
Data compression continues to evolve, with traditional information theory methods being widely used for compressing text, images, and videos. Recently, there has been growing interest in leveraging Generative AI for predictive compression techniques. This paper 1 introduces a lossless text compression approach using a…
Emir Öztürk, Altan Mesut, Stefano Cirillo
Learning-based data compression methods have gained significant attention in recent years. Although these methods achieve higher compression ratios compared to traditional techniques, their slow processing times make them less suitable for compressing large datasets, and they are generally more effective for short…
M Baritha Begum, N. Deepa, Mueen Uddin, Rajesh Kaluri + 2 more
'Maha Abdelhaq' 'Raed Alsaqour'] Data stored on physical storage devices and transmitted over communication channels often have a lot of redundant information, which can be reduced through compression techniques to conserve space and reduce the time it takes to transmit the data. The need for adequate security…
Fajia Sun, Long Qian
DNA has been pursued as a compelling medium for digital data storage during the past decade. While large-scale data storage and random access have been achieved in artificial DNA, the synthesis cost keeps hindering DNA data storage from popularizing into daily life. In this study, we proposed a more efficient paradigm…
Anas Al-okaily, Abdelghani Tbakhi
Data compression is a challenging and increasingly important problem. As the amount of data generated daily continues to increase, efficient transmission and storage have never been more critical. In this study, a novel encoding algorithm is proposed, motivated by the compression of DNA data and associated…
Sebastian Schmidt, Shahbaz Khan, Jarno Alanko, Alexandru I. Tomescu
Kmer-based methods are widely used in bioinformatics, which raises the question of what is the smallest practically usable representation (i.e. plain text) of a set of kmers. We propose a polynomial algorithm computing a minimum such representation (which was previously posed as a potentially NP-hard open problem), as…
Subhankar Roy, Dilip Kumar Maity, Anirban Mukhopadhyay
Deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) sequence compressors for novel species frequently face challenges when processing wide-scale raw, FASTA, or multi-FASTA structured data. For years, molecular sequence databases have favored the widely used general-purpose Gzip and Zstd compressors. The absence of…
Pétur Helgi Einarsson, Páll Melsted
We describe a compression scheme for BUS files and an implementation of the algorithm in the bustools software. Our compression algorithm yields smaller file sizes than gzip, at significantly faster compression and decompression speeds. We evaluated our algorithm on 533 BUS files from scRNA-seq experiments with a total…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
Foad Nazari, Sneh Patel, Melissa LaRocca, Ryan Czarny + 2 more
As sequencing becomes more accessible, there is an acute need for novel compression methods to efficiently store this data. Omics technologies can enhance biomedical research and individualize patient care, but they demand immense storage capabilities, especially when applied to longitudinal studies. Addressing the…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…
Luke Staniscia, Yun William Yu
Because of the rapid generation of data, the study of compression algorithms to reduce storage and transmission costs is important to bioinformaticians. Much of the focus has been on sequence data, including both genomes and protein amino acid sequences stored in FASTA files. Current standard practice is to use an…