24 papers · ranked by Valyu relevance
Logan S. Whitehouse, Dylan Ray, Daniel R. Schrider
As population genetics data increases in size new methods have been developed to store genetic information in efficient ways, such as tree sequences. These data structures are computationally and storage efficient, but are not interchangeable with existing data structures used for many population genetic inference…
Logan S Whitehouse, Dylan D Ray, Daniel R Schrider, Diogo Meyer
As population genetic data increase in size, new methods have been developed to store genetic information in efficient ways, such as tree sequences. These data structures are computationally and storage efficient but are not interchangeable with existing data structures used for many population genetic inference…
Michel Rigo, Manon Stipulanti
The nth term of an automatic sequence is the output of a deterministic finite automaton fed with the representation of n in a suitable numeration system. In this paper, instead of considering automatic sequences built on a numeration system with a regular numeration language, we consider those built on languages…
Jacob Gilbert, Chih Hao Wu, Marina Knittel, Alejandro A. Schäffer + 2 more
Understanding and comparing tumor evolutionary histories is fundamental to cancer genomics. Clonal trees, used to model tumor progression, are rooted, unordered trees in which each node represents a subclone labeled by a set of distinct mutations. To compare two clonal trees, we introduce omlta, the optimal multi-label…
Nathan Fox
Labeled infinite trees provide combinatorial interpretations for many integer sequences generated by nested recurrence relations. Typically, such sequences are monotone increasing. Several of these sequences also have straightforward descriptions in terms of how often each value in the sequence occurs. In this paper…
Toufik Mansour, Gökhan Yıldırım
We introduce an algorithmic approach based on generating tree method for enumerating the inversion sequences with various pattern-avoidance restrictions. For a given set of patterns, we propose an algorithm that outputs either an accurate description of the succession rules of the corresponding generating tree or an…
Julia A. Palacios, Anand Bhaskar, Filippo Disanto, Noah A. Rosenberg
Evolutionary models used for describing molecular sequence variation suppose that at a non-recombining genomic segment, sequences share ancestry that can be represented as a genealogy-a rooted, binary, timed tree, with tips corresponding to individual sequences. Under the infinitely-many-sites mutation model, mutations…
Joe Sawada, James W. Sears, A. Trautrim, Aaron Williams
Classic cycle-joining techniques have found widespread application in creating universal cycles for a diverse range of combinatorial objects, such as shorthand permutations, weak orders, orientable sequences, and various subsets of k-ary strings, including de Bruijn sequences. In the most favorable scenarios, these…
Lukas Hübner, Alexandros Stamatakis
The field of population genetics attempts to advance our understanding of evolutionary processes. It has applications, for example, in medical research, wildlife conservation, and – in conjunction with recent advances in ancient DNA sequencing technology – studying human migration patterns over the past few thousand…
Youngjun Park, Juhyeon Kim, Joonsuk Huh
A guide tree directs the sequence alignment order for multiple sequence alignment (MSA). Distance-based guide trees are widely used in MSA tools. However, many traditional algorithms for constructing these guide trees are unscalable and require a time complexity of O(N^2^) to O(N^3^), where N is the number of input…
Lars Arvestad
Distance-based methods for inferring evolutionary trees are important subroutines in computational biology, sometimes as a first step in a statistically more robust phylogenetic method. The most popular method is Neighbor Joining, mainly to to its relatively good accuracy, but Neighbor Joining has a cubic time…
Kou Hamada, Sankardeep Chakraborty, Seungbum Jo, Takuto Koriyama + 2 more
and Efficient Implementation of Average-Case Optimal RMQs Authors: ['Kou Hamada' 'Sankardeep Chakraborty' 'Seungbum Jo' 'Takuto Koriyama' 'Kunihiko Sadakane' 'Srinivasa Rao Satti'] Tree covering is a technique for decomposing a tree into smaller-sized trees with desirable properties, and has been employed in various…
Andrew C. Riley, Daniel A. Ashlock, Steffen P. Graether, Ruriko Yoshida
'Ruriko Yoshida'] Intrinsically disordered proteins (IDPs) are proteins that lack a stable 3D structure but maintain a biological function. It has been frequently suggested that IDPs are difficult to align because they tend to have fewer conserved residues compared to ordered proteins, but to our knowledge this has…
Wei Wei, David Koslicki
Distance-guided tree construction with unknown tree topology and branch lengths has been a long studied problem. In contrast, distance-guided branch lengths assignment with fixed tree topology has not yet been systematically investigated, despite having significant applications. In this paper, we provide a formal…
Ting Wang, Zu-Guo Yu, Jinyan Li
Traditional alignment-based methods meet serious challenges in genome sequence comparison and phylogeny reconstruction due to their high computational complexity. Here, we propose a new alignment-free method to analyze the phylogenetic relationships (classification) among species. In our method, the dynamical language…
Joel Gustafsson, Peter Norberg, Jan R. Qvick-Wester, Alexander Schliep
'Alexander Schliep'] Background Alignment-free methods are a popular approach for comparing biological sequences, including complete genomes. The methods range from probability distributions of sequence composition to first and higher-order Markov chains, where a k-th order Markov chain over DNA has $4^k$ formal…
Julia A. Palacios, Anand Bhaskar, Filippo Disanto, Noah A. Rosenberg
Evolutionary models used for describing molecular sequence variation suppose that at a non-recombining genomic segment, sequences share ancestry that can be represented as a genealogy—a rooted, binary, timed tree, with tips corresponding to individual sequences. Under the infinitely-many-sites mutation model, mutations…
Liang Liu, Lili Yu, Shaoyuan Wu, Jonathan Arnold + 3 more
'Christopher Whalen' 'Charles Davis' 'Scott Edwards'] Accurate reconstruction of species trees often relies on the quality of input gene trees estimated from molecular sequences. Previous studies suggested that if the sequence length is fixed, the maximum likelihood may produce biased gene trees which subsequently…
Yuanyuan Qi, Mohammed El-Kebir
Cancer phylogenies are key to understanding tumor evolution. There exists many important downstream analyses that takes as input a single or small number of trees. However, due to uncertainty, one typically infers many, equally-plausible phylogenies from bulk DNA sequencing data of tumors. We introduce Sapling, a…
Authors not listed
RNA molecules fold into complex three-dimensional structures that determine their function. A wide range of mathematical frameworks, such as chord diagrams, fatgraphs, and context-free grammars, have been used to represent these structures; however, these models have largely been developed from mathematical motivations…
Jonas Schaub, Julian Zander, Achim Zielesny, Christoph Steinbeck
The concept of molecular scaffolds as defining core structures of organic molecules is utilised in many areas of chemistry and cheminformatics, e.g. drug design, chemical classification, or the analysis of high-throughput screening data. Here, we present Scaffold Generator, a comprehensive open library for the…
Vladimir Kondratyev, Marian Dryzhakov, Timur Gimadiev, Dmitriy Slutskiy
In this work, we provide further development of the junction tree variational autoencoder (JT VAE) architecture in terms of implementation and application of the internal feature space of the model. Pretraining of JT VAE on a large dataset and further optimization with a regression model led to a latent space that can…
Richard Apodaca
Despite its widespread use, Simplified Molecular Input Line Entry System (SMILES) remains underspecified. The lack of a detailed specification encourages improvisation by software developers, complicates data standardization efforts, and undermines extension development. Balsa, a reformulation of SMILES, addresses…
Authors not listed
We present a simple yet efficient random (brute-force) algorithm for constructing solvated molecular systems. By placing solvent molecules at random positions and orientations within a simulation box, we circumvent the complexities typically associated with more sophisticated packing algorithms. The main computational…