Search · four archives
Search · four archives
24 papers · ranked by Valyu relevance
Emily G. Light, Morgan E. Prior, Noah M. Daniels, Najib Ishaq
Motivation: The multiple sequence alignment (MSA) problem has been extensively studied, with numerous approaches developed over recent years. With the rapid growth of sequence data, there is an increasing need for fast and accurate MSA tools that scale effectively to large datasets. Building on our previous work on…
Yanming Wei, Zhaoyang Huang, Pinglu Zhang, Yizheng Wang + 3 more
Multiple sequence alignment (MSA) is a fundamental problem in bioinformatics. The quality of sequence alignment significantly impacts biological sequence analysis, especially that in next-generation sequencing . MSA results are widely used in various applications, including de novo genome assembly , detection of…
Krishnendu Sinha
Multiple sequence alignment (MSA) underpins comparative genomics, evolutionary analysis, and structural inference. Despite extensive methodological development, most widely used alignment algorithms rely on static amino acid substitution matrices that encode global average substitution tendencies and are inherently…
Yixiao Zhai, Zitong Zhang, Zhen Li
The reliability of multiple sequence alignment (MSA) results directly determines the credibility of the conclusions drawn from biological research. However, MSA is inherently an NP-hard problem, making it theoretically impossible to guarantee a globally optimal solution. Consequently, in addition to developing more…
Nimrod Serok, Ksenia Polonsky, Haim Ashkenazy, Itay Mayrose + 2 more
Multiple sequence alignment (MSA) inference is a central task in molecular evolution and comparative genomics, and the reliability of downstream analyses, including phylogenetic inference, depends critically on alignment quality. Despite this importance, most widely used MSA methods optimize the sum-of-pairs (SP)…
Krishnendu Sinha, Pier Luigi Martelli
Multiple sequence alignment (MSA) is a foundational operation in computational biology, underpinning phylogenetic reconstruction, functional motif discovery, evolutionary rate estimation, and comparative structural analysis (, ). Despite decades of algorithmic development, accurate alignment of divergent protein…
Minh Hoang, Isabel Armour-Garb, Mona Singh
Multiple sequence alignment (MSA) is a foundational task in computational biology, under-pinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning…
Haodong Liu, Pinglu Zhang, Yanming Wei, Qinzhong Tian + 3 more
Partial order alignment (POA) has emerged as a fundamental component in long-read error correction, assembly and pangenomics. However, conventional POA algorithms are limited by high time and memory requirements, making them inefficient for large-scale datasets. Here, we present minipoa, a fast and memory-efficient POA…
Yoshiki Kanazawa, Naphan Benchasattabuse, Michal Hajdušek, Rodney Van Meter
We formulate a structure-informed multiple sequence alignment problem, denoted MSA-S. The model abstracts biological sequences as strings and structural information as designated position-pairs. It augments a fixed pairwise string score, defined by a fixed non-gap symbol-pair scoring rule and fixed affine gap…
A. Burak Gulhan, Richard Burhans, Robert Harris, Mahmut Kandemir + 2 more
Advances in sequencing and assembly allow the creation of thousands of genome assemblies. However, producing multiple alignments required for their analysis lags behind due to the time-consuming process of pairwise alignment, typically performed by the slow but sensitive tool lastZ. Here, we develop KegAlign, an…
Jannik Olbrich, Enno Ohlebusch
Modern genomic analyses increasingly rely on pangenomes, that is, representations of the genome of entire populations. The simplest representation of a pangenome is a set of individual genome sequences. Compared to e.g. sequence graphs, this has the advantage that efficient exact search via indexes based on the…
Miguel Graça, Aleksandar Ilic
State-of-the-art multiple sequence alignment (MSA) algorithms are based on progressive approaches that rely on pairwise sequence alignment (PSA) to generate guide trees to align all sequences. Given an evidenced explosion in genomic data availability, research efforts have focused on accelerating PSA on…
Mattis Bodynek, Lucía Martín-Fernández, Julia Haag, Ben Bettisworth + 1 more
Multiple Sequence Alignment (MSA) constitutes an important and frequent operation in molecular sequence data analysis. There exist numerous tools, algorithms, and criteria to infer an MSA. This plethora of available approaches to MSA may induced an ensemble of divergent MSAs for the same underlying unaligned sequence…
Hao Xuan, Hongyang Sun, Xiangtao Liu, Hanyuan Zhang + 2 more
Sequence alignment underpins nearly every facet of modern genomics, from genetic testing and cancer profiling to functional genome annotation. Yet, despite decades of algorithmic innovation, most existing aligners remain narrowly optimized for specific tasks, fragmenting analytical workflows and limiting…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
Background The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated…
Zhezhen Yu, Dan Levy
Aligning sequencing reads to short tandem repeats (STRs) is challenging: the number of repeat copies in a read often differs from the reference, and small changes inside and around the repeat can lead to many competing alignments. We introduce NW-flex, a simple extension of classical sequence alignment that addresses…
Valentin Gorgodian, Olivier Poirot, Alain Schmitt, Virginie Collomb + 4 more
Phylogenetic analysis has become a standard approach across many areas of biology, yet the growing complexity of phylogenetic methods and software remains a major obstacle for non-specialists. Since its launch in 2008, Phylogeny.fr has provided an accessible web platform for building phylogenetic trees using widely…
Simeng Zhang, Xinying Liu, Jun Lou, Mudi Jiang + 2 more
—The rapid development of high-throughput sequencing technologies has led to an explosive increase in biological sequence data, making sequence clustering a fundamental task in large-scale bioinformatics analyses. Unlike traditional clustering problems, biological sequence clustering faces unique challenges due to the…
Nadja Nolte, Marko Petek, Pablo Angulo Lara, Logan Mulroney + 3 more
Inaccurate allele and gene expression counts due to map bias and genome ambiguity lead to high false positive and false negative rates in studies of allelic imbalance. We demonstrate that long read RNA sequencing (RNA-seq) and straightforward quality control measures can be used to reduce bias in allele counts in case…
Authors not listed
Predicting how chemical modifications affect drug binding is central to rational drug design. Free Energy Perturbation (FEP) calculations provide accurate estimates of these binding affinity changes, but existing methods often require substantial computational resources and expert knowledge. Here we present QligFEP…
Authors not listed
RNA molecules fold into complex three-dimensional structures that determine their function. A wide range of mathematical frameworks, such as chord diagrams, fatgraphs, and context-free grammars, have been used to represent these structures; however, these models have largely been developed from mathematical motivations…
Authors not listed
Machine learning models are increasingly applied to heterogeneous materials datasets spanning different synthesis routes, measurement protocols, and structural classes. Although multi-task and representation-learning approaches are commonly used to improve predictive performance, the latent representations learned by…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Authors not listed
Accurately measuring compound binding affinities is key to driving the pharmaceutical development process. Rigorous physics-based in silico approaches, particularly alchemical free energy methods, have become a gold standard tool for estimating compound affinity changes. Here we present the results of a large-scale…