8 papers · ranked by Valyu relevance
Lena Collienne, Harry Richman, David H. Rich, Mary Barker + 2 more
Deep learning offers hope for more efficient phylogenetic inference methods. However, it has yet to have the transformative effect on phylogenetics that it has had in other fields. Here we present a novel approach that combines deep learning with concepts behind current successful phylogenetic algorithms. Specifically…
Rohit Fenn, Amit Fenn
Any process that generates information at a constant rate into a branching hierarchy must embed into hyperbolic space: exponentially growing lineages cannot pack into polynomial-growth Euclidean geometry. We derive a geometric state equation, κ = (h ln 2/(n−1))^2^, relating the curvature κ of the embedding manifold to…
Edo Dotan, Asaf Schers, Elya Wygoda, Tal Pupko + 1 more
Accurate inference of phylogenetic trees is fundamental to evolutionary biology, yet existing methods rely on complex pipelines involving multiple sequence alignment, explicit evolutionary models, and computationally intensive tree search procedures. Here, we present BetaInfer, a generative framework that reformulates…
Niklas Müller, H. Steven Scholte, Iris I. A. Groen
In real-world vision, the human brain needs to process large amounts of information to effectively interact with its environment. It is well established that our visual system has specialized regions to process information efficiently, such as scene-, face-, and object-selective areas, which can be uncovered using…
Marvin De los Santos
In large-scale phylogenetic analysis, incorporating translation awareness is critical to account for the genotypic and phenotypic dimensions underlying biological diversification. Covary is a machine learning-based framework that analyzes, clusters, and compares genetic sequences through alignment-free…
Jacob Gilbert, Chih Hao Wu, Marina Knittel, Alejandro A. Schäffer + 2 more
Understanding and comparing tumor evolutionary histories is fundamental to cancer genomics. Clonal trees, used to model tumor progression, are rooted, unordered trees in which each node represents a subclone labeled by a set of distinct mutations. To compare two clonal trees, we introduce omlta, the optimal multi-label…
Adrian Tkachenko, Sepehr Salem, Ayotomiwa Ezekiel Adeniyi, Zülal Bingöl + 6 more
High-throughput sequencing (HTS) enables population-scale genomics but generates massive datasets, creating bottlenecks in storage, transfer, and analysis. FASTQ, the standard format for over two decades, stores one byte per base and one byte per quality score, leading to inefficient I/O, high storage costs, and…
Matthieu Vilain, Stéphane Aris-Brosou
The ever-growing amount of available biological data leads modern analysis to be performed on large datasets. Unfortunately, bioinformatics tools for preprocessing and analyzing data are not always designed to treat such large amounts of data efficiently. Notably, this is the case when encoding DNA and RNA sequences…