12 papers · ranked by Valyu relevance
Kshitij Tayal, Naveen Sivadasan, Rajgopal Srinivasan
We consider the computational problem of phasing an individual genotype sample given a collection of known haplotypes in the population. We give a fast and accurate algorithm GPhase for reconstructing haplotype pair consistent with input genotype. It uses the coalescent based mutation model of Stephens and Donnelly…
Juanjo Bermúdez
Genome assembly is a fundamental tool for biological research. Particularly, in microbiology, where budgets per sample are often scarce, it can make the difference between an inconclusive result and a fully valid conclusion. Identifying new strains or estimating the relative abundance of quasi-species in a sample are…
Diego Darriba, David Posada
Several strategies have been proposed to assign substitution models in phylogenomic datasets, or partitioning. The accuracy of these methods, and most importantly, their impact on phylogenetic estimation has not been thoroughly assessed using computer simulations. We simulated multiple partitioning scenarios to…
Yukun Yang, Wolfgang Maass
Most current methods for goal-directed action selection in the face of changing goals and contingencies require DNNs or LLMs. Therefore they are less suited for implementation in edge devices, where low energy-consumption is imperative. The brain shows that similar functionality can be produced with just 20W, even with…
Benjamin T. James, Brian B. Luczak, Hani Z. Girgis
Sequence clustering is a fundamental step in analyzing DNA sequences. Widely-used software tools for sequence clustering utilize greedy approaches that are not guaranteed to produce the best results. These tools are sensitive to one parameter that determines the similarity among sequences in a cluster. Often times, a…
Stavros I. Dimitriadis, Eirini Messaritaki, Derek K. Jones
The human brain is a complex network of volumes of tissue (nodes) that are interconnected by white matter tracts (edges). It can be represented as a graph to allow us to use graph theory to gain insight into normal human development and brain disorders. Most graph theoretical metrics measure either whole-network…
Arseny Shur, Ido Tziony, Yaron Orenstein
Minimizers are sampling schemes which are ubiquitous in almost any high-throughput sequencing analysis. Assuming a fixed alphabet of size σ, a minimizer is defined by two positive integers k, w and a linear order ρ on k-mers. A sequence is processed by a sliding window algorithm that chooses in each window of length w…
Mattia Eluchans, Gian Luca Lancia, Antonella Maselli, Marco D’Alessando + 2 more
We humans are capable of solving challenging planning problems, but the range of adaptive strategies that we use to address them are not yet fully characterized. Here, we designed a series of problem-solving tasks that require planning at different depths. After systematically comparing the performance of participants…
Kuan-Hao Chao, Pei-Wei Chen, Sanjit A. Seshia, Ben Langmead
A Wheeler graph represents a collection of strings in a way that is particularly easy to index and query. Such a graph is a practical choice for representing a graph-shaped pangenome, and it is the foundation for current graph-based pangenome indexes. However, there are no practical tools to visualize or to check…
Ragnar Groot Koerkamp, Igor Martayan
Because of the rapidly-growing amount of sequencing data, computing sketches of large textual datasets has become an essential preprocessing task. These sketches are typically much smaller than the input sequences, but preserve sufficient information for downstream analysis. Minimizers are an especially popular…
Wilfried Agbeto, Camille Coti, Vladimir Reinharz
Subgraph isomorphism is a combinatorial problem that involves finding one or all occurrences of a pattern graph within a target graph. Subgraph isomorphism has numerous applications in fields such as biology, chemistry, social network analysis, and pattern recognition. Although subgraph isomorphism is generally…
Mikko Rautiainen, Veli Mäkinen, Tobias Marschall
Graphs are commonly used to represent sets of sequences. Either edges or nodes can be labeled by sequences, so that each path in the graph spells a concatenated sequence. Examples include graphs to represent genome assemblies, such as string graphs and de Bruijn graphs, and graphs to represent a pan-genome and hence…