21 papers · ranked by Valyu relevance
Vu H. Nguyen, Hien T. Nguyen, Hieu N. Duong, Vaclav Snasel
We propose an efficient method for compressing Vietnamese text using n-gram dictionaries. It has a significant compression ratio in comparison with those of state-of-the-art methods on the same dataset. Given a text, first, the proposed method splits it into n-grams and then encodes them based on n-gram dictionaries.…
Chowdhury Mofizur Rahman, Mahbub E. Sobhani, Anika Tasnim Rodela, Swakkhar Shatabda
'Swakkhar Shatabda'] Abstract—Text compression shrinks textual data while keeping crucial information, eradicating constraints on storage, bandwidth, and computational efficacy. The integration of lossless compression techniques with transformer-based text decompression has received negligible attention, despite the…
Beniamin Stecuła, Kinga Stecuła, Adrian Kapczyński, Antonio Puliafito
'Antonio Puliafito'] The goal of the research was to study the possibility of using the planned language Esperanto for text compression, and to compare the results of the text compression in Esperanto with the compression in natural languages, represented by Polish and English. The authors performed text compression in…
Mohammad Hosseini
—Today, with the growing demands of information storage and data transfer, data compression is becoming increasingly important. Data Compression is a technique which is used to decrease the size of data. This is very useful when some huge files have to be transferred over networks or being stored on a data storage…
Swathi Shree Narashiman, Nitin Chandrachoodan
Data compression continues to evolve, with traditional information theory methods being widely used for compressing text, images, and videos. Recently, there has been growing interest in leveraging Generative AI for predictive compression techniques. This paper 1 introduces a lossless text compression approach using a…
Emir Öztürk, Altan Mesut, Stefano Cirillo
Learning-based data compression methods have gained significant attention in recent years. Although these methods achieve higher compression ratios compared to traditional techniques, their slow processing times make them less suitable for compressing large datasets, and they are generally more effective for short…
Juncai Xu, Weidong Zhang, Qingwen Ren, Xin Xie + 1 more
There is a special type of text which the order of the rows makes no difference (e.g., a word list). To compress these special texts, the traditional lossless compression method is not the ideal choice. A new method that can achieve better compression results for this type of texts is proposed. The texts are…
Rajneil Baruah, Vaskar Deka, M. P. Bhuyan
With the rapid growing of data and number of applications, there is a crucial need of dictionary based reversible transformation techniques to increase the efficiency of the compression algorithms and hence contribute towards the enhancement in compression ratio. Performance analysis of compression methods in…
Anas Al-okaily, Abdelghani Tbakhi
Data compression is a challenging and increasingly important problem. As the amount of data generated daily continues to increase, efficient transmission and storage has never been more critical. In this study, a novel encoding algorithm is proposed, motivated by the compression of DNA data and associated…
M Baritha Begum, N. Deepa, Mueen Uddin, Rajesh Kaluri + 2 more
'Maha Abdelhaq' 'Raed Alsaqour'] Data stored on physical storage devices and transmitted over communication channels often have a lot of redundant information, which can be reduced through compression techniques to conserve space and reduce the time it takes to transmit the data. The need for adequate security…
Justin Kim, Rahul Varki, Marco Oliva, Christina Boucher
The RePair compression algorithm produces a context-free grammar by iteratively substituting the most frequently occurring pair of consecutive symbols with a new symbol until all consecutive pairs of symbols appear only once in the compressed text. It is widely used in the settings of bioinformatics, machine learning…
Fajia Sun, Long Qian
DNA has been pursued as a compelling medium for digital data storage during the past decade. While large-scale data storage and random access have been achieved in artificial DNA, the synthesis cost keeps hindering DNA data storage from popularizing into daily life. In this study, we proposed a more efficient paradigm…
Anas Al-okaily, Abdelghani Tbakhi
Data compression is a challenging and increasingly important problem. As the amount of data generated daily continues to increase, efficient transmission and storage have never been more critical. In this study, a novel encoding algorithm is proposed, motivated by the compression of DNA data and associated…
Rahul Varki, Travis Gagie, Christina Boucher
Among grammar-based compression techniques, RePair is a notable offline encoding scheme known for its simplicity and powerful combinatorial properties, producing compact grammars by repeatedly replacing the most frequent adjacent pairs of symbols, known as bigrams. However, RePair’s memory usage scales poorly with…
I Made
— People tend to store a lot of files inside theirs storage. When the storage nears it limit, they then try to reduce those files size to minimum by using data compression software. In this paper we propose a new algorithm for data compression, called jbit encoding (JBE). This algorithm will manipulates each bit of…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
J.M. Lázaro-Guevara, K.M. Garrido
Undeveloped countries like Guatemala, where access to high-speed internet connections is limited, downloading and sharing Biological information of thousands of Mega Bits is a huge problem for the beginning and development of Bioinformatics. Based on that information is an urgent necessity to find a better way to share…
David G. Nagy, Balázs Török, Gergő Orbán
It has extensively been documented that human memory exhibits a wide range of systematic distortions, which have been associated with resource constraints. Resource constraints on memory can be formalised in the normative framework of lossy compression, however traditional lossy compression algorithms result in…
Felix Zeller, Chieh-Min Hsieh, Wilke Dononelli, Tim Neudecker
The field of liquid-phase and solid-state high-pressure chemistry has exploded since the advent of the diamond anvil cell, an experimental technique that allows the application of pressures up to several hundred gigapascal. To complement high-pressure experiments, a large number of computational tools have been…
Authors not listed
Cost effective and reliable hydrogen compression remains a challenging barrier in the wide-spread adoption of hydrogen as an energy carrier. The prevailing technology of mechanical compression suffers from several drawbacks, some of which can be addressed by non-mechanical compression strategies (e.g., electrochemical…