10 papers · ranked by Valyu relevance
Kexin Niu, Maxat Kulmanov, Robert Hoehndorf
Current machine learning methods for enzyme function prediction primarily treat proteins as independent entities, ignoring the metabolic context in which they operate. This reductionist approach often generates biologically implausible annotations that fail to satisfy stoichiometric or thermodynamic constraints. While…
Jing Xie, Qi Duan
Biological pathway analysis often requires identifying interventions that block reachability to an undesirable state, such as a disease-associated module, toxic byproduct, or adverse phenotype, while preserving reachability among essential biological functions. Motivated by this setting, we study the Reachability…
Hugo Magalhães, Jonas Weber, Gunnar W. Klau, Tobias Marschall + 1 more
Variation of sequence copy number (CN) between individuals can be associated with phenotypical differences. Consequently, CN calling is an important step for disease association and identification, as well as for genome assembly validation. Traditionally, CN calling is done by mapping sequencing reads to a linear…
Frans Zdyb, Julius B. Kirkegaard
Biological image and video analysis is full of discrete decisions: whether an object is present, which multi-hypothesis detections are real, whether two detections are tracking the same object, or whether a cell divides or not. Standard pipelines resolve these locally and in sequence, e.g through non-max suppression…
Mahsa Faizrahnemoon, Jens Luebeck, King L. Hung, Suhas Rao + 7 more
Extrachromosomal DNA (ecDNA) plays a key role in cancer pathology. EcDNAs mediate high oncogene amplification and expression and worse patient outcomes. Accurately determining the structure of these circular molecules is essential for understanding their function, yet reconstructing ecDNA cycles from sequencing data…
Ke Chen, Abhishek Talesara, Sanchal Thakkar, Mingfu Shao
The minimum flow decomposition problem abstracts a set of key tasks in bioinformatics, including metagenome and transcriptome assembly. These tasks, collectively known as multi-assembly, aim to reconstruct multiple genomic sequences from reads obtained from mixed samples. The reads are first organized into a directed…
Xinyu Gu, Stefan Ivanovic, Daniel W. Feng, Mohammed El-Kebir
Summarizing a collection P of related RNA secondary structures is a key challenge in applications like evolutionary analysis, alternative fold studies and mRNA vaccine design. This requires both clustering the input structures into similar groups and identifying the core structural motifs on which they agree or differ.…
Joseph Guhlin, Peter Dearden
Genomic best linear unbiased prediction (GBLUP) is widely used for genomic selection in livestock and crop breeding. There is growing interest in connecting machine learning with genomic breeding value prediction. Although the BLUP formula is itself mathematically differentiable, existing implementations do not expose…
Ghanshyam Chandra, William T. Doan, Daniel Gibney
Genotyping is the task of identifying the genetic variants present in a sample from sequencing data, and it is a fundamental problem in computational biology. Existing genotyping approaches typically rely on either haplotype reference panels or pangenome graphs. Compared to haplotype reference panels, pangenome graphs…
A.J.R. Cotter
A simulator, ‘ECOLPS’ in R, is developed and trialed for ecological studies of closed aquatic ecosystems. Its constraint-based approach contrasts with function-based models widely applied in ecology. Total gross production (ΣGP) by ‘wild components’ (= species/life stages, grouped by ecological roles) is maximized…