24 papers · ranked by Valyu relevance
Noam Teyssier, Alexander Dobin
Modern genomics produces billions of sequencing records per run, which are typically stored as gzip-compressed FASTQ files. While this format is widely used, it is not optimal for high-throughput processing due to its reliance on single-threaded decompression and sequential parsing of irregularly sized records. This…
Enrique Canessa
A comprehensive study of the properties of finite (0,1) binary systems from the mathematical viewpoint of quantum theory is presented. This is a quantum-inspired extension of the GenomeBits model to characterize observed genome sequences, where a complex wavefunction $\psi _{n}$ is considered as an analogous…
Juan Carlos Nuño, Francisco J. Muñoz, Carlo Cattani
We study some properties of binary sequences generated by random substitutions of constant length. Specifically, assuming the alphabet ${0,1}$, we consider the following asymmetric substitution rule of length k: $0\rightarrow〈0,0,…,0〉$ and $1\rightarrow〈Y_{1},Y_{2},…,Y_{k}〉$, where $Y_{i}$ is a Bernoulli random…
Tuvi Etzion
—Binary self-dual sequences have been considered and analyzed throughout the years, and they have been used for various applications. Motivated by a construction for single-track Gray codes, we examine the structure and recursive constructions for binary and non-binary self-dual sequences. The feedback shift registers…
Jin-yan Hu, Gang Yan, Tao Wang
The study of various living complex systems by system identification method is important, and the identification of the problem is even more challenging when dealing with a dynamic nonlinear system of discrete time. A well-established model based on kernel functions for input of the maximum length sequence (m-sequence)…
Bernard Costa, Marcus V. C. Baldo, Carolina Feher da Silva
Adaptive human behaviour depends on the ability to detect regularities and probabilistic structures within a noisy environment. Repeated binary choice tasks, in which individuals predict one of two possible outcomes, have long served as a fundamental tool for investigating learning, reward processing, and…
Authors not listed
Sequence is the critical determinant of macromolecular function, yet current polymer design approaches often optimize monomer composition and ratios while ignoring sequence. This creates poorly defined design spaces for active learning that miss the vast combinatorial landscape of sequence possibilities. We introduce…
Miguel Beltrá, Sara D. Cardell, Verónica Requena
The binary binomial sequences correspond to the diagonals of the Pascal's triangle modulo 2. They have interesting properties such as they form a basis of the linear space of all binary sequences with period a power of 2. Other properties of these sequences (period, linear complexity, construction rules or relations…
Aldo C. Martínez, Aldo Solis, Rafael Díaz Hernández Rojas, Alfred B. U’Ren + 2 more
'Alfred B. U’Ren' 'Jorge G. Hirsch' 'Isaac Pérez Castillo'] Pseudo-random number generators are widely used in many branches of science, mainly in applications related to Monte Carlo methods, although they are deterministic in design and, therefore, unsuitable for tackling fundamental problems in security and…
Daniel Katz
Pseudorandom sequences are used extensively in communications and remote sensing. Correlation provides one measure of pseudorandomness, and low correlation is an important factor determining the performance of digital sequences in applications. We consider the problem of constructing pairs (f, g) of sequences such that…
Grenville J. Croll, Boris Ryabko
The distribution of prime numbers has long been viewed as a balance between order and randomness. In this work, we investigate the relationship between entropy, periodicity, and primality through the computational framework of the binary derivative. We prove that periodic numbers are composite in all bases except for a…
Ian Holmes
We describe a strategy for constructing codes for DNA-based information storage by serial composition of weighted finite-state transducers. The resulting state machines can integrate correction of substitution errors; synchronization by interleaving watermark and periodic marker signals; conversion from binary to…
Eric Chen, Adam Ge, Andrew Kalashnikov, Tanya Khovanova + 7 more
'Evin Liang' 'Mira Lubashev' 'Matthew Qian' 'Rohith Raghavan' 'Benjamin Taycher' 'Samuel J. Wang'] In this paper, we generalize a lot of facts from John Conway and Alex Ryba's paper, The extra Fibonacci series and the Empire State Building, where we replace the Fibonacci sequence with the Tribonacci sequence. We study…
ARIYA SHAJII, IBRAHIM NUMANAGIĆ, RIYADH BAGHDADI, BONNIE BERGER + 1 more
'SAMAN AMARASINGHE' ''] The scope and scale of biological data are increasing at an exponential rate, as technologies like next-generation sequencing are becoming radically cheaper and more prevalent. Over the last two decades, the cost of sequencing a genome has dropped from $100 million to nearly $100-a factor of…
Samuele Girotto, Matteo Comin, Cinzia Pizzi
Background Spaced-seeds, i.e. patterns in which some fixed positions are allowed to be wild-cards, play a crucial role in several bioinformatics applications involving substrings counting and indexing, by often providing better sensitivity with respect to k-mers based approaches. K-mers based approaches are usually…
David L Abel, Jack T Trevors
Genetic algorithms instruct sophisticated biological organization. Three qualitative kinds of sequence complexity exist: random (RSC), ordered (OSC), and functional (FSC). FSC alone provides algorithmic instruction. Random and Ordered Sequence Complexities lie at opposite ends of the same bi-directional sequence…
Bartosz Sobolewski, Maciej Ulas
In this note we investigate the solutions of certain meta-Fibonacci recurrences of the form f(n) = f(n − f(n − 1)) + f(n − 2) for various sets of initial conditions. In the case when f(n) = 1 for n ≤ 1, we prove that the resulting integer sequence is closely related to the function counting binary partitions of a…
Daniel B. Shapiro
Set h0i! = 1, and note that n 0 f = n n f = 1. Also, n k f is undefined when k > n and when n or k is negative. (Some authors set n k f = 0 when k > n.) That factorial formula helps explain the symmetry:
Ben Chen, Richard H. Chen, Joshua Guo, Tanya Khovanova + 7 more
'Neil Malur' 'Nastia Polina' 'Poonam Sahoo' 'Anuj Sakarda' 'Nathan C. Sheffield' 'Armaan Tipirneni'] We discuss properties of integers in base 3/2. We also introduce many new sequences related to base 3/2. Some sequences discuss patterns related to integers in base 3/2. Other sequence are analogues of famous base-10…
Roland Wittler
To index or compare sequences efficiently, often k-mers, i.e., substrings of fixed length k, are used. For efficient indexing or storage, k-mers are often encoded as integers, e.g., applying some bijective mapping between all possible σ^k^ k-mers and the interval [0, σ^k^ −1], where σ is the alphabet size. In many…
N. J. A. Sloane
Until 1973 there was no database of integer sequences. Someone coming across the sequence 1, 2, 4, 9, 21, 51, 127, . . . would have had no way of discovering that it had been studied since 1870 (today these are called the Motzkin numbers, and form entry A001006 in the database). Everything changed in 1973 with the…
Babu Bassa
In this communication the author describes a software tool named "ChameleonSort". The software program, developed by the present author is useful in the sorting of biological sequence variants like those accumulating mutations while diverging from the common ancestors. Examples include viral protein variants, protein…
Patrick Kunzmann
Alignment searches are fast heuristic methods to identify similar regions between two sequences. This group of algorithms is ubiquitously used in a myriad of software to find homologous sequences or to map sequence reads to genomes. Often the first step in alignment searches is k-mer decomposition: listing all…
Nicholas Tjahjono, Evgeni Penev, Boris Yakobson
Programmable self-assembly provides a promising avenue to improve upon traditional synthesis and create multi-component materials with emergent properties and arbitrary nanoscale complexity. However, its most successful realizations utilizing DNA often use complicated arduous procedures that result in low yields. Here…