8 papers · ranked by Valyu relevance
Kenny Daily, Paul Rigor, Scott Christley, Xiaohui Xie + 1 more
Background High-throughput sequencing (HTS) technologies play important roles in the life sciences by allowing the rapid parallel sequencing of very large numbers of relatively short nucleotide sequences, in applications ranging from genome sequencing and resequencing to digital microarrays and ChIP-Seq experiments. As…
Francesca Console, Giuseppe D’Aquanno, Giuseppe Antonio Di Luna, Leonardo Querzoni + 1 more
'Leonardo Querzoni' 'Aswani Kumar Cherukuri'] In this article we propose the first multi-task benchmark for evaluating the performances of machine learning models that work on low level assembly functions. While the use of multi-task benchmark is a standard in the natural language processing (NLP) field, such practice…
Kun Tu, Dariusz Puchala, Jun Chen, Sadaf Salehkalaibar
In this paper, we address the problem of m-gram entropy variable-to-variable coding, extending the classical Huffman algorithm to the case of coding m-element (i.e., m-grams) sequences of symbols taken from the stream of input data for $m>1$. We propose a procedure to enable the determination of the frequencies of the…
Zhi Ping, Dongzhao Ma, Xiaoluo Huang, Shihong Chen + 4 more
'Fei Guo' 'Sha Joe Zhu' 'Yue Shen'] Title: Abstract The information explosion has led to a rapid increase in the amount of data requiring physical storage. However, in the near future, existing storage methods (i.e., magnetic and optical media) will be insufficient to store these exponentially growing data. Therefore…
Nithin Nagaraj, Arun Somani
Error detection is a fundamental need in most computer networks and communication systems in order to combat the effect of noise. Error detection techniques have also been incorporated with lossless data compression algorithms for transmission across communication networks. In this paper, we propose to incorporate a…
Jesús E. Garca, Verónica A. González-López, Gustavo H. Tasca, Karina Y. Yaginuma + 1 more
In the framework of coding theory, under the assumption of a Markov process $(X_{t})$ on a finite alphabet $A,$ the compressed representation of the data will be composed of a description of the model used to code the data and the encoded data. Given the model, the Huffman’s algorithm is optimal for the number of bits…
Katharina Mir, Klaus Neuhaus, Martin Bossert, Steffen Schober + 1 more
'Eshel Ben-Jacob'] We consider the design and evaluation of short barcodes, with a length between six and eight nucleotides, used for parallel sequencing on platforms where substitution errors dominate. Such codes should have not only good error correction properties but also the code words should fulfil certain…
David Kracht, Steffen Schober
Background Barcode multiplexing is a key strategy for sharing the rising capacity of next-generation sequencing devices: Synthetic DNA tags, called barcodes, are attached to natural DNA fragments within the library preparation procedure. Different libraries, can individually be labeled with barcodes for a joint…