25 papers · ranked by Valyu relevance
Yusuke Imai
We create a sequence version of calculus. First, we define equivalence, some fundamental operations, differential, and integral for sequences. Then, we propose sequence versions of identity function, power function, exponential function, hyperbolic function, trigonometric function, and also find sequence versions of…
Ariya Shajii, Ibrahim Numanagić, Alexander T. Leighton, Haley Greenyer + 2 more
Exponentially-growing next-generation sequencing data requires high-performance tools and algorithms. Nevertheless, the implementation of high-performance computational genomics software is inaccessible to many scientists because it requires extensive knowledge of low-level software optimization techniques, forcing…
Andreas Döring, David Weese, Tobias Rausch, Knut Reinert
Background The use of novel algorithmic techniques is pivotal to many important problems in life science. For example the sequencing of the human genome [1] would not have been possible without advanced assembly algorithms. However, owing to the high speed of technological progress and the urgent need for…
Ovidiu Popa, Ellen Oldenburg, Oliver Ebenhöh
Today massive amounts of sequenced metagenomic and metatranscriptomic data from different ecological niches and environmental locations are available. Scientific progress depends critically on methods that allow extracting useful information from the various types of sequence data. Here, we will first discuss types of…
N. J. A. Sloane
Until 1973 there was no database of integer sequences. Someone coming across the sequence 1, 2, 4, 9, 21, 51, 127, . . . would have had no way of discovering that it had been studied since 1870 (today these are called the Motzkin numbers, and form entry A001006 in the database). Everything changed in 1973 with the…
Wensheng Gan, Jerry Chun‐Wei Lin, Jiexiong Zhang, Han‐Chieh Chao + 2 more
'Hamido Fujita' 'Philip S. Yu'] Utility is an important concept in economics. A variety of applications consider utility in real-life situations, which has lead to the emergence of utility-oriented mining (also called utility mining) in the recent decade. Utility mining has attracted a great amount of attention, but…
Aleksey Buzmakov, Elias Egho, Nicolas Jay, Sergei O. Kuznetsov + 2 more
'Amedeo Napoli' 'Chedy Raïssi'] Nowadays data sets are available in very complex and heterogeneous ways. Mining of such data collections is essential to support many realworld applications ranging from healthcare to marketing. In this work, we focus on the analysis of "complex" sequential data by means of interesting…
Babu Bassa
In this communication the author describes a software tool named "ChameleonSort". The software program, developed by the present author is useful in the sorting of biological sequence variants like those accumulating mutations while diverging from the common ancestors. Examples include viral protein variants, protein…
Diana M. Müssgens, Fredrik Ullén
Transfer (i.e., the application of a learned skill in a novel context) is an important and desirable outcome of motor skill learning. While much research has been devoted to understanding transfer of explicit skills the mechanisms of skill transfer after incidental learning remain poorly understood. The aim of this…
Authors not listed
Sequence is the critical determinant of macromolecular function, yet current polymer design approaches often optimize monomer composition and ratios while ignoring sequence. This creates poorly defined design spaces for active learning that miss the vast combinatorial landscape of sequence possibilities. We introduce…
Chun Liang, Feng Sun, Haiming Wang, Junfeng Qu + 3 more
'Lee H Pratt' 'Marie-Michèle Cordonnier-Pratt'] Background Processing raw DNA sequence data is an especially challenging task for relatively small laboratories and core facilities that produce as many as 5000 or more DNA sequences per week from multiple projects in widely differing species. To meet this challenge, we…
K.S Kong, E.Y.K Ng
The work showed that the integrated suite of software tools for detecting criminals using DNA databases has achieved the overall objective by providing a working platform for sequence analysis. The work also demonstrated that by integrating BLAST and FASTA (two widely used and freely available algorithms), plus an…
ARIYA SHAJII, IBRAHIM NUMANAGIĆ, RIYADH BAGHDADI, BONNIE BERGER + 1 more
'SAMAN AMARASINGHE' ''] The scope and scale of biological data are increasing at an exponential rate, as technologies like next-generation sequencing are becoming radically cheaper and more prevalent. Over the last two decades, the cost of sequencing a genome has dropped from $100 million to nearly $100-a factor of…
Ke Chen, Vinamratha Pattar, Mingfu Shao
Sequence similarity estimation is essential for many bioinformatics tasks, including functional annotation, phylogenetic analysis, and overlap graph construction. Alignment-free methods aim to solve large-scale sequence similarity estimation by mapping sequences to more easily comparable features that can approximate…
Theresa B. Loveless, Courtney K. Carlson, Vincent J. Hu, Catalina A. Dentzel Helmy + 4 more
Genetically encoded DNA recorders noninvasively convert transient biological events into durable mutations in a cell’s genome, allowing for the later reconstruction of cellular experiences using high-throughput DNA sequencing^1^. Existing DNA recorders have achieved high-information recording^2–14^, durable…
Authors not listed
Cyclic peptides become attractive therapeutic candidates due to their diverse biological activities. However, existing deep learning-based sequence design models, such as ProteinMPNN, are primarily optimized using cross-entropy loss and often overlook the unique topological constraints of cyclic peptides. This limits…
David Novák, Petr Volný, Pavel Zezula
Subsequence matching has appeared to be an ideal approach for solving many problems related to the fields of data mining and similarity retrieval. It has been shown that almost any data class (audio, image, biometrics, signals) is or can be represented by some kind of time series or string of symbols, which can be seen…
Declan Devlin, Korbinian Moeller, Iro Xenidou-Dervou, Bert Reynvoet + 1 more
'Francesco Sella'] Both adults and children are slower at judging the ordinality of non-consecutive sequences (e.g., 1-3-5) than consecutive sequences (e.g., 1-2-3). It has been suggested that the processing of non-consecutive sequences is slower because it conflicts with the intuition that only count-list sequences…
Haohan Zhu, George Kollios, Vassilis Athitsos
This paper proposes a general framework for matching similar subsequences in both time series and string databases. The matching results are pairs of query subsequences and database subsequences. The framework finds all possible pairs of similar subsequences if the distance measure satisfies the "consistency" property…
Authors not listed
This paper presents a simplified model of iterative compound optimization in drug/agrochemical discovery. Compounds are represented as binary strings, with project evolution simulated through random bit changes. The model reproduces key statistical features of real projects, including activity distributions and…
Jasper Linthorst, Marc Hulsman, Henne Holstege, Marcel Reinders
The emergence of third generation sequencing technologies has brought near perfect de-novo genome assembly within reach. This clears the way towards reference-free detection of genomic variations. In this paper, we introduce a novel concept for aligning whole-genomes which allows the alignment of multiple genomes.…
Guillaume Marçais, C.S. Elder, Carl Kingsford
Sequences equivalent to their reverse complements (i.e., double-stranded DNA) have no analogue in text analysis and non-biological string algorithms. Despite this striking difference, algorithms designed for computational biology (e.g., sketching algorithms) are designed and tested in the same way as classical string…
Kazuyoshi Tsuchiya, Chiaki Ogawa, Yasuyuki Nogami, Satoshi Uehara
Pseudorandom number generators are required to generate pseudorandom numbers which have good statistical properties as well as unpredictability in cryptography. An m-sequence is a linear feedback shift register sequence with maximal period over a finite field. M-sequences have good statistical properties, however we…
Valeriy Titarenko, Sofya Titarenko
Technical progress in computer hardware made it possible to access and process large amounts of data even on budget workstations. Therefore new or existing alignment algorithms may use large index files to increase performance. Spaced seeds with large weights reduce the number of possible locations of a read within a…
Fuzhan Rahmanian, Robert M. Lee, Dominik Linzner, Kathrin Michel + 4 more
Predicting and monitoring battery life early and across chemistries is a significant challenge due to the plethora of degradation paths, form factors, and electrochemical testing protocols. Existing models typically translate poorly across different electrode, electrolyte, and additive materials, mostly require a fixed…