24 papers · ranked by Valyu relevance
Yusuke Imai
We create a sequence version of calculus. First, we define equivalence, some fundamental operations, differential, and integral for sequences. Then, we propose sequence versions of identity function, power function, exponential function, hyperbolic function, trigonometric function, and also find sequence versions of…
N. J. A. Sloane
Until 1973 there was no database of integer sequences. Someone coming across the sequence 1, 2, 4, 9, 21, 51, 127, . . . would have had no way of discovering that it had been studied since 1870 (today these are called the Motzkin numbers, and form entry A001006 in the database). Everything changed in 1973 with the…
László Mérai, Arne Winterhof
Many automatic sequences, such as the Thue-Morse sequence or the Rudin-Shapiro sequence, have some desirable features of pseudorandomness such as a large linear complexity and a small well-distribution measure. However, they also have some disastrous properties in view of certain applications. For example, the majority…
Theresa B. Loveless, Courtney K. Carlson, Vincent J. Hu, Catalina A. Dentzel Helmy + 4 more
Genetically encoded DNA recorders noninvasively convert transient biological events into durable mutations in a cell’s genome, allowing for the later reconstruction of cellular experiences using high-throughput DNA sequencing^1^. Existing DNA recorders have achieved high-information recording^2–14^, durable…
Babu Bassa
In this communication the author describes a software tool named "ChameleonSort". The software program, developed by the present author is useful in the sorting of biological sequence variants like those accumulating mutations while diverging from the common ancestors. Examples include viral protein variants, protein…
Hao Xuan, Hongyang Sun, Xiangtao Liu, Hanyuan Zhang + 2 more
Sequence alignment underpins nearly every facet of modern genomics, from genetic testing and cancer profiling to functional genome annotation. Yet, despite decades of algorithmic innovation, most existing aligners remain narrowly optimized for specific tasks, fragmenting analytical workflows and limiting…
Authors not listed
Sequence is the critical determinant of macromolecular function, yet current polymer design approaches often optimize monomer composition and ratios while ignoring sequence. This creates poorly defined design spaces for active learning that miss the vast combinatorial landscape of sequence possibilities. We introduce…
Dana G. Korssjoen, Biyao Li, Stefan Steinerberger, Raghavendra Tripathi + 1 more
'Raghavendra Tripathi' 'Ruimin Zhang'] Abstract. We investigate a method of generating a graph G = (V, E) out of an ordered list of n distinct real numbers a1, . . . , an. These graphs can be used to test for the presence of combinatorial structure in the sequence. We describe sequences exhibiting intricate hidden…
Ke Chen, Vinamratha Pattar, Mingfu Shao
Sequence similarity estimation is essential for many bioinformatics tasks, including functional annotation, phylogenetic analysis, and overlap graph construction. Alignment-free methods aim to solve large-scale sequence similarity estimation by mapping sequences to more easily comparable features that can approximate…
Donald R Campbell Jr, Timothee Cezard, Sveinung Gundersen, Andrew D. Yates + 8 more
Reference genomes are foundational to genomics but suffer from widespread ambiguity and incompatibility due to inconsistent naming, undocumented differences, and lack of formal mechanisms for comparison. To address this, we introduce the GA4GH refget Sequence Collections (seqcol) standard. Refget seqcol is a framework…
Chris J. Mitchell, Peter Wild
This paper describes new, simple, recursive methods of construction for orientable sequences, i.e. periodic binary sequences in which any n-tuple occurs at most once in a period in either direction. As has been previously described, such sequences have potential applications in automatic position-location systems…
Tan Li, Mengshan Li, Yan Wu, Yelin Li + 1 more
The efficient analysis and interpretation of biological sequence data remain major challenges in bioinformatics. Graphical representation, as an emerging and effective visualization technique, offers a more intuitive method for analyzing DNA sequences. However, many visualization approaches are dispersed across…
Zhu, Jian, Lin, Zhidong + 8 more
—Discovering valuable insights from rich data is a crucial task for exploratory data analysis. Sequential pattern mining (SPM) has found widespread applications across various domains. In recent years, low-utility sequential pattern mining (LUSPM) has shown strong potential in applications such as intrusion detection…
Xiaoshu Ma, Yusha Wang, Ruikai Jia, Hua Ye
The Single Molecule Real Time (SMRT) system developed by Pacific Biosciences applies the principle of synthesis while sequencing and uses the SMRT chip as the sequencing carrier. The high starting amount and good integrity of DNA required by PacBio library construction has always been a headache. Generally, the total…
Authors not listed
Cyclic peptides become attractive therapeutic candidates due to their diverse biological activities. However, existing deep learning-based sequence design models, such as ProteinMPNN, are primarily optimized using cross-entropy loss and often overlook the unique topological constraints of cyclic peptides. This limits…
Declan Devlin, Korbinian Moeller, Iro Xenidou-Dervou, Bert Reynvoet + 1 more
'Francesco Sella'] Number order processing is thought to be characterised by a reverse distance effect whereby consecutive sequences (e.g., 1-2-3) are processed faster than non-consecutive sequences (e.g., 1-3-5). However, there is accumulating evidence that the reverse distance effect is not consistently observed. In…
Wenxiong Zhou, Li Kang, Shuo Qiao, Haifeng Duan + 19 more
High-throughput sequencing technologies generate a vast number of DNA sequence reads simultaneously, which are subsequently analyzed using the information contained within these fragmented reads. The assessment of sequencing technology relies on information efficiency, which measures the amount of information entropy…
Declan Devlin, Korbinian Moeller, Iro Xenidou-Dervou, Bert Reynvoet + 1 more
'Francesco Sella'] Both adults and children are slower at judging the ordinality of non-consecutive sequences (e.g., 1-3-5) than consecutive sequences (e.g., 1-2-3). It has been suggested that the processing of non-consecutive sequences is slower because it conflicts with the intuition that only count-list sequences…
Authors not listed
This paper presents a simplified model of iterative compound optimization in drug/agrochemical discovery. Compounds are represented as binary strings, with project evolution simulated through random bit changes. The model reproduces key statistical features of real projects, including activity distributions and…
Guillaume Marçais, C.S. Elder, Carl Kingsford
Sequences equivalent to their reverse complements (i.e., double-stranded DNA) have no analogue in text analysis and non-biological string algorithms. Despite this striking difference, algorithms designed for computational biology (e.g., sketching algorithms) are designed and tested in the same way as classical string…
Pedro A. Bernaola-Galván, Pedro Carpena, Cristina Gómez-Martín, José L. Oliver + 2 more
The concept of a genome signature broadly refers to characteristic patterns in DNA sequences that enable the identification and comparison of species or individuals, often without requiring sequence alignment. Such signatures have applications ranging from forensic identification of individuals to cancer genomics. In…
Robin D. P. Zhou
Inspired by the definition of modified ascent sequences, we introduce a new class of integer sequences called revised ascent sequences. These sequences are defined as Cayley permutations where each entry is a leftmost occurrence if and only if it serves as an ascent bottom. We construct a bijection between ascent…
J. P. Correia
This paper investigates the complexity of DNA sequences in maize and soybean using the multifractal detrended fluctuation analysis (MF-DFA) method, chaos game representation (CGR), and the complexity-entropy plane approach. The study aims to understand the patterns and structures of these DNA sequences, which can…
Fuzhan Rahmanian, Robert M. Lee, Dominik Linzner, Kathrin Michel + 4 more
Predicting and monitoring battery life early and across chemistries is a significant challenge due to the plethora of degradation paths, form factors, and electrochemical testing protocols. Existing models typically translate poorly across different electrode, electrolyte, and additive materials, mostly require a fixed…