20 papers · ranked by Valyu relevance
Paridhi Latawa, Nuh Aydın
| 1 | Abstract | | 2 | | --- | --- | --- | --- | | 2 | | Introduction | 2 | | 3 | | Convolutional Codes | 3 | | | 3.1 | Encoding of Binary Convolutional Codes | 3 | | | 3.2 | Decoding Convolutional Codes | 11 | | | 3.3 | Truncated Viterbi Decoding | 13 | | 4 | | DNA Codes | 17 | | | 4.1 | Constraints for the…
Kenny Daily, Paul Rigor, Scott Christley, Xiaohui Xie + 1 more
Background High-throughput sequencing (HTS) technologies play important roles in the life sciences by allowing the rapid parallel sequencing of very large numbers of relatively short nucleotide sequences, in applications ranging from genome sequencing and resequencing to digital microarrays and ChIP-Seq experiments. As…
H. Yamamoto, Masato Tsuchihashi, Junya Honda
We propose almost instantaneous fixed-to-variable-length (AIFV) codes such that two (resp. K − 1) code trees are used if code symbols are binary (resp. K-ary for K ≥ 3), and source symbols are assigned to incomplete internal nodes in addition to leaves. Although the AIFV codes are not instantaneous codes, they are…
Ian Holmes
We describe a strategy for constructing codes for DNA-based information storage by serial composition of weighted finite-state transducers. The resulting state machines can integrate correction of substitution errors; synchronization by interleaving watermark and periodic marker signals; conversion from binary to…
Inbal Preuss, Michael Rosenberg, Zohar Yakhini, Leon Anavy
With the world generating digital data at an exponential rate, DNA has emerged as a promising archival medium. It offers a more efficient and long-lasting digital storage solution due to its durability, physical density, and high information capacity. Research in the field includes the development of encoding schemes…
Francesca Console, Giuseppe D’Aquanno, Giuseppe Antonio Di Luna, Leonardo Querzoni + 1 more
'Leonardo Querzoni' 'Aswani Kumar Cherukuri'] In this article we propose the first multi-task benchmark for evaluating the performances of machine learning models that work on low level assembly functions. While the use of multi-task benchmark is a standard in the natural language processing (NLP) field, such practice…
Neha Periwal, Priya Sharma, Pooja Arora, Saurabh Pandey + 2 more
Classification among coding (CDS) and non-coding RNA (ncRNA) sequences is a challenge and several machine learning models have been developed for the same. Since the frequency of curated coding sequences is many-folds as compared to that of the ncRNAs, we devised a novel approach to work with the complete datasets from…
А. В. Анисимов, Igor O. Zavadskyi
Variable-length splittable codes are derived from encoding sequences of ordered integer pairs, where one of the pair's components is upper bounded by some constant, and the other one is any positive integer. Each pair is encoded by the concatenation of two fixed independent prefix encoding functions applied to the…
Ahmed Hareedy, Beyza Dabak, Robert Calderbank
Constrained codes are used to prevent errors from occurring in various data storage and data transmission systems. They can help in increasing the storage density of magnetic storage devices, in managing the lifetime of electronic storage devices, and in increasing the reliability of data transmission over wires. Over…
Kun Tu, Dariusz Puchala, Jun Chen, Sadaf Salehkalaibar
In this paper, we address the problem of m-gram entropy variable-to-variable coding, extending the classical Huffman algorithm to the case of coding m-element (i.e., m-grams) sequences of symbols taken from the stream of input data for $m>1$. We propose a procedure to enable the determination of the frequencies of the…
Xuyang Zhao, Junyao Li, Qingyuan Fan, Jing Dai + 5 more
DNA, as the origin for the genetic information flow, has also been a compelling alternative to non-volatile information storage medium. Reading digital information from this highly dense but lightweighted medium nowadays relied on conventional next-generation sequencing (NGS), which involves ‘wash and read’ cycles for…
Zhi Ping, Dongzhao Ma, Xiaoluo Huang, Shihong Chen + 4 more
'Fei Guo' 'Sha Joe Zhu' 'Yue Shen'] Title: Abstract The information explosion has led to a rapid increase in the amount of data requiring physical storage. However, in the near future, existing storage methods (i.e., magnetic and optical media) will be insufficient to store these exponentially growing data. Therefore…
Nithin Nagaraj, Arun Somani
Error detection is a fundamental need in most computer networks and communication systems in order to combat the effect of noise. Error detection techniques have also been incorporated with lossless data compression algorithms for transmission across communication networks. In this paper, we propose to incorporate a…
Jesús E. Garca, Verónica A. González-López, Gustavo H. Tasca, Karina Y. Yaginuma + 1 more
In the framework of coding theory, under the assumption of a Markov process $(X_{t})$ on a finite alphabet $A,$ the compressed representation of the data will be composed of a description of the model used to code the data and the encoded data. Given the model, the Huffman’s algorithm is optimal for the number of bits…
Subhash Kak
Mathematically, ternary coding is more efficient than binary coding. It is little used in computation because technology for binary processing is already established and the implementation of ternary coding is more complicated, but remains relevant in algorithms that use decision trees and in communications. In this…
Robert Bamler
Entropy coding is the backbone data compression. Novel machine-learning based compression methods often use a new entropy coder called Asymmetric Numeral Systems (ANS) [Duda et al., 2015], which provides very close to optimal bitrates and simplifies [Townsend et al., 2019] advanced compression techniques such as…
Rod Rinkus
The brain is believed to implement probabilistic reasoning and to represent information via population, or distributed, coding. Most previous population-based probabilistic (PPC) theories share several basic properties: 1) continuous-valued neurons (units); 2) fully/densely-distributed codes, i.e., all/most coding…
Amir Said
Entropy coding, compression, complexity This introduction to arithmetic coding is divided in two parts. The first explains how and why arithmetic coding works. We start presenting it in very general terms, so that its simplicity is not lost under layers of implementation details. Next, we show some of its basic…
Katharina Mir, Klaus Neuhaus, Martin Bossert, Steffen Schober + 1 more
'Eshel Ben-Jacob'] We consider the design and evaluation of short barcodes, with a length between six and eight nucleotides, used for parallel sequencing on platforms where substitution errors dominate. Such codes should have not only good error correction properties but also the code words should fulfil certain…
David Kracht, Steffen Schober
Background Barcode multiplexing is a key strategy for sharing the rising capacity of next-generation sequencing devices: Synthetic DNA tags, called barcodes, are attached to natural DNA fragments within the library preparation procedure. Different libraries, can individually be labeled with barcodes for a joint…