21 papers · ranked by Valyu relevance
Rahman, Md. Atiqur, Rabbi, MM Fazle
The rapid growth of digital data has heightened the demand for efficient lossless compression methods. However, existing algorithms exhibit trade-offs: some achieve high compression ratios, others excel in encoding or decoding speed, and none consistently perform best across all dimensions. This mismatch complicates…
Yibo Yang, Stephan Mandt, Lucas Theis
Neural compression is the application of neural networks and other machine learning methods to data compression. Recent advances in statistical machine learning have opened up new possibilities for data compression, allowing compression algorithms to be learned end-to-end from data using powerful generative models such…
Mohammad Hosseini
—Today, with the growing demands of information storage and data transfer, data compression is becoming increasingly important. Data Compression is a technique which is used to decrease the size of data. This is very useful when some huge files have to be transferred over networks or being stored on a data storage…
Kun Tu, Dariusz Puchala, Jun Chen, Sadaf Salehkalaibar
In this paper, we address the problem of m-gram entropy variable-to-variable coding, extending the classical Huffman algorithm to the case of coding m-element (i.e., m-grams) sequences of symbols taken from the stream of input data for $m>1$. We propose a procedure to enable the determination of the frequencies of the…
Tomasz Krokosz, Jarogniew Rykowski, Małgorzata Zajęcka, Robert Brzoza-Woch + 2 more
'Robert Brzoza-Woch' 'Leszek Rutkowski' 'Amitabh Mishra'] Modern, commonly used cryptosystems based on encryption keys require that the length of the stream of encrypted data is approximately the length of the key or longer. In practice, this approach unnecessarily complicates strong encryption of very short messages…
Ziguang Li, Chao Huang, Xuliang Wang, Haibo Hu + 6 more
'Dongbo Bu' 'Quan Yu' 'Wen Gao' 'Xingwu Liu' 'Ming Li'] The LLMs may be seen to approximate the uncomputable Solomonoff induction. Therefore, under this new uncomputable paradigm, we present LMCompress. LMCompress shatters all previous lossless compression algorithms, doubling the lossless compression ratios of JPEG-XL…
Sagnik Banerjee, Carson Andorf
Advancement in technology has enabled sequencing machines to produce vast amounts of genetic data, causing an increase in storage demands. Most genomic software utilizes read alignments for several purposes including transcriptome assembly and gene count estimation. Herein we present, ABRIDGE, a state-of-the-art…
Lei M. Li
We consider the lossless compression bound of any individual data sequence. Conceptually, its Kolmogorov complexity is such a bound yet uncomputable. The Shannon source coding theorem states that the average compression bound is nH, where n is the number of words and H is the entropy of an oracle probability…
Marina Galchenkova, Alexandra Tolstikova, Bjarne Klopprogge, Janina Sprenger + 7 more
Various approaches for lossless and lossy compression are evaluated, and suitable quality assessment metrics for serial crystallographic data - used in combination with lossy data reduction - are described.
Anas Al-okaily, Abdelghani Tbakhi
Data compression is a challenging and increasingly important problem. As the amount of data generated daily continues to increase, efficient transmission and storage have never been more critical. In this study, a novel encoding algorithm is proposed, motivated by the compression of DNA data and associated…
Amal Altamimi, Belgacem Ben Youssef, Oleg Sergiyenko, Wendy Flores-Fuentes + 2 more
'Wendy Flores-Fuentes' 'Julio Cesar Rodríguez-Quiñonez' 'Jesús Elías Miranda-Vega'] Rapid and continuous advancements in remote sensing technology have resulted in finer resolutions and higher acquisition rates of hyperspectral images (HSIs). These developments have triggered a need for new processing techniques…
Kavindu Jayasooriya, Sasha P. Jenner, Pasindu Marasinghe, Udith Senanayake + 5 more
Nanopore sequencing is an increasingly central tool for genomics. Despite rapid advances in the field, large data volumes and computational bottlenecks continue to pose major challenges. Here we introduce ex-zd, a new data compression strategy that helps address the large size of raw signal data generated during…
David Podgorelec, Damjan Strnad, Ivana Kolingerová, Borut Žalik + 1 more
'Jun Chen'] After a boom that coincided with the advent of the internet, digital cameras, digital video and audio storage and playback devices, the research on data compression has rested on its laurels for a quarter of a century. Domain-dependent lossy algorithms of the time, such as JPEG, AVC, MP3 and others…
Herbert J. Bernstein, Alexei S. Soares, Kimberly Horvat, Jean Jakoncic
New higher-count-rate, integrating, large area X-ray detectors with framing rates as high as 17,400 images per second are beginning to be available. These will soon be used for specialized MX experiments but will require optimal lossy compression algorithms to enable systems to keep up with data throughput. Some…
M Baritha Begum, N. Deepa, Mueen Uddin, Rajesh Kaluri + 2 more
'Maha Abdelhaq' 'Raed Alsaqour'] Data stored on physical storage devices and transmitted over communication channels often have a lot of redundant information, which can be reduced through compression techniques to conserve space and reduce the time it takes to transmit the data. The need for adequate security…
Luke Staniscia, Yun William Yu
Because of the rapid generation of data, the study of compression algorithms to reduce storage and transmission costs is important to bioinformaticians. Much of the focus has been on sequence data, including both genomes and protein amino acid sequences stored in FASTA files. Current standard practice is to use an…
Junjie Tong, Miaoshan Lu, Bichen Peng, Shaowei An + 2 more
The size of high-resolution mass spectrometry (HRMS) data has been increasing significantly. Several lossy compressors have been developed for higher compression rate. Currently, a comprehensive evaluation of what and how MS data (m/z and intensities) with precision losses would affect data processing is absent.…
Vasileios Alevizos, Nikitas Gerolimos, Sabrina Edralin, C. J. Xu + 4 more
'Akebu Simasiku' 'Georgios Priniotakis' 'George A. Papakostas' 'Zongliang Yue'] Abstract—One requirement of maintaining digital information is storage. With the latest advances in the digital world, new emerging media types have required even more storage space to be kept than before. In fact, in many cases it is…
Bin Duan, Logan A Walker, Bin Xie, Wei Jie Lee + 3 more
Recent advances in microscopy have pushed imaging data generation to an unprecedented scale. While scientists benefit from higher spatiotemporal resolutions and larger imaging volumes, the increasing data size presents significant storage, visualization, sharing, and analysis challenges. Lossless compression typically…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…