10 papers · ranked by Valyu relevance
Noam Teyssier, Alexander Dobin
Modern genomics produces billions of sequencing records per run, which are typically stored as gzip-compressed FASTQ files. While this format is widely used, it is not optimal for high-throughput processing due to its reliance on single-threaded decompression and sequential parsing of irregularly sized records. This…
Samuel Planton, Fosca Al Roumi, Liping Wang, Stanislas Dehaene
According to the language of thought hypothesis, regular sequences are compressed in human working memory using recursive loops akin to a mental program that predicts future items. We tested this theory by probing working memory for 16-item sequences made of two sounds. We recorded brain activity with functional MRI…
Bernard Costa, Marcus V. C. Baldo, Carolina Feher da Silva
Adaptive human behaviour depends on the ability to detect regularities and probabilistic structures within a noisy environment. Repeated binary choice tasks, in which individuals predict one of two possible outcomes, have long served as a fundamental tool for investigating learning, reward processing, and…
Yusei Kobori, Satoshi Mizuta
Graphical representation of DNA sequences is one of the most popular techniques of alignment-free sequence comparison. In this article, we propose a new method for extracting features of DNA sequences represented by binary images, in which we estimate the similarity between DNA sequences by the frequency histograms of…
Ian Holmes
We describe a strategy for constructing codes for DNA-based information storage by serial composition of weighted finite-state transducers. The resulting state machines can integrate correction of substitution errors; synchronization by interleaving watermark and periodic marker signals; conversion from binary to…
Valeriy Titarenko, Sofya Titarenko
Technical progress in computer hardware made it possible to access and process large amounts of data even on budget workstations. Therefore new or existing alignment algorithms may use large index files to increase performance. Spaced seeds with large weights reduce the number of possible locations of a read within a…
Wenxiong Zhou, Li Kang, Shuo Qiao, Haifeng Duan + 19 more
High-throughput sequencing technologies generate a vast number of DNA sequence reads simultaneously, which are subsequently analyzed using the information contained within these fragmented reads. The assessment of sequencing technology relies on information efficiency, which measures the amount of information entropy…
Patrick Kunzmann
Alignment searches are fast heuristic methods to identify similar regions between two sequences. This group of algorithms is ubiquitously used in a myriad of software to find homologous sequences or to map sequence reads to genomes. Often the first step in alignment searches is k-mer decomposition: listing all…
Mehmet Ali Tibatan, Mustafa Sarisaman
We investigate the quantum behavior encountered in palindromes within DNA structure. In particular, we reveal the unitary structure of usual palindromic sequences found in genomic DNAs of all living organisms, using the Schwinger’s approach. We clearly demonstrate the role played by palindromic configurations with…
Roland Wittler
To index or compare sequences efficiently, often k-mers, i.e., substrings of fixed length k, are used. For efficient indexing or storage, k-mers are often encoded as integers, e.g., applying some bijective mapping between all possible σ^k^ k-mers and the interval [0, σ^k^ −1], where σ is the alphabet size. In many…