14 papers · ranked by Valyu relevance
Yohei M. Rosen, Benedict J. Paten
Hidden Markov models of haplotype inheritance such as the Li and Stephens model allow for computationally tractable probability calculations using the forward algorithms as long as the representative reference panel used in the model is sufficiently small. Specifically, the monoploid Li and Stephens model and its…
Fabio F. de Oliveira, Leonardo A. Dias, Marcelo A. C. Fernandes
In bioinformatics, alignment is an essential technique for finding similarities between biological sequences. Usually, the alignment is performed with the Smith-Waterman (SW) algorithm, a well-known sequence alignment technique of high-level precision based on dynamic programming. However, given the massive data volume…
Tung Dang, Hirohisa Kishino
Random forest (RF) captures complex feature patterns that differentiate groups of samples and is rapidly being adopted in microbiome studies. However, a major challenge is the high dimensionality of microbiome datasets. They include thousands of species or molecular functions of particular biological interest. This…
Jonathan Terhorst
I present phlash, a new Bayesian method for inferring population history from whole genome sequence data. phlash is population history learning by averaging sampled histories: it works by drawing random, low-dimensional projections of the coalescent intensity function from the posterior distribution of a psmc-like…
David S. Lawrie
Forward Wright-Fisher simulations are powerful in their ability to model complex demography and selection scenarios, but suffer from slow execution on the CPU, thus limiting their usefulness. The single-locus Wright-Fisher forward algorithm is, however, exceedingly parallelizable, with many steps which are so-called…
Lukas Hecker, Amita Giri, Dimitrios Pantazis, Amir Adler
Magnetoencephalography (MEG) and electroencephalography (EEG) are widely employed techniques for the in-vivo measurement of neural activity with exceptional temporal resolution. Modeling the neural sources underlying these signals is of high interest for both neuroscience research and pathology. The method of…
Tung Dang, Alan S. R. Fermin, Maro G. Machizawa
Neuroimaging data is complex and high-dimensional that poses challenges for machine learning (ML) applications. Of varieties of reasons contributing on accuracy decoding, variable feature selection is one of crucial steps for determining target feature in data analysis, especially in the context of neuroimaging studies…
Regev Schweiger, Yaniv Erlich, Shai Carmi
Hidden Markov models (HMMs) are powerful tools for modeling processes along the genome. In a standard genomic HMM, observations are drawn, at each genomic position, from a distribution whose parameters depend on a hidden state; the hidden states evolve along the genome as a Markov chain. Often, the hidden state is the…
Eric Alcaide, Stella Biderman, Amalio Telenti, M. Cyrus Maher
The conversion of proteins between internal and cartesian coordinates is a limiting step in many pipelines, such as molecular dynamics simulations and machine learning models. This conversion is typically carried out by sequential or parallel applications of the Natural extension of Reference Frame (NeRF) algorithm.…
Paola Bonizzoni, Christina Boucher, Davide Cozzi, Travis Gagie + 3 more
The positional Burrows–Wheeler Transform (PBWT) was presented in 2014 by Durbin as a means to find all maximal haplotype matches in h sequences containing w variation sites in 𝒪(hw)-time. This time complexity of finding maximal haplotype matches using the PBWT is a significant improvement over the naïve…
Francesco Donnarumma, Thomas Parr, Karl Friston, James Whittington + 1 more
How the brain plans and maintains sequences of future actions remains a central question in systems neuroscience. Recent studies in the frontal cortex have revealed that multiple elements of a sequence are represented simultaneously in separable neural subspaces, challenging classical serial models of sequential…
Jordan M. Eizenga, Benedict Paten
Modern genomic sequencing data is trending toward longer sequences with higher accuracy. Many analyses using these data will center on alignments, but classical exact alignment algorithms are infeasible for long sequences. The recently proposed WFA algorithm demonstrated how to perform exact alignment for long, similar…
Beren Millidge, Mufeng Tang, Mahyar Osanlouy, Nicol S. Harper + 1 more
One of the key problems the brain faces is inferring the state of the world from a sequence of dynamically changing stimuli, and it is not yet clear how the sensory system achieves this task. A well-established computational framework for describing perceptual processes in the brain is provided by the theory of…
Mikko Rautiainen, Veli Mäkinen, Tobias Marschall
Graphs are commonly used to represent sets of sequences. Either edges or nodes can be labeled by sequences, so that each path in the graph spells a concatenated sequence. Examples include graphs to represent genome assemblies, such as string graphs and de Bruijn graphs, and graphs to represent a pan-genome and hence…