24 papers · ranked by Valyu relevance
Hao Xuan, Hongyang Sun, Xiangtao Liu, Hanyuan Zhang + 2 more
Sequence alignment underpins nearly every facet of modern genomics, from genetic testing and cancer profiling to functional genome annotation. Yet, despite decades of algorithmic innovation, most existing aligners remain narrowly optimized for specific tasks, fragmenting analytical workflows and limiting…
Nimrod Serok, Ksenia Polonsky, Haim Ashkenazy, Itay Mayrose + 2 more
Multiple sequence alignment (MSA) inference is a central task in molecular evolution and comparative genomics, and the reliability of downstream analyses, including phylogenetic inference, depends critically on alignment quality. Despite this importance, most widely used MSA methods optimize the sum-of-pairs (SP)…
Miguel Graça, Aleksandar Ilic
State-of-the-art multiple sequence alignment (MSA) algorithms are based on progressive approaches that rely on pairwise sequence alignment (PSA) to generate guide trees to align all sequences. Given an evidenced explosion in genomic data availability, research efforts have focused on accelerating PSA on…
Hasitha Kaushan, Santiago Marco-Sola, Erik Garrison, Pjotr Prins + 1 more
Storing millions of sequence alignments from large-scale genomic comparisons requires efficient compression methods. While fixed-size alignment encodings offer uniform spacing and bounded reconstruction cost, they cannot adapt to variable alignment complexity across sequences, missing compression opportunities in…
Minh Hoang, Isabel Armour-Garb, Mona Singh
Multiple sequence alignment (MSA) is a foundational task in computational biology, under-pinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning…
Yanming Wei, Zhaoyang Huang, Pinglu Zhang, Yizheng Wang + 3 more
Multiple sequence alignment (MSA) is a fundamental problem in bioinformatics. The quality of sequence alignment significantly impacts biological sequence analysis, especially that in next-generation sequencing . MSA results are widely used in various applications, including de novo genome assembly , detection of…
Nimrod Serok, Ksenia Polonsky, Haim Ashkenazy, Itay Mayrose + 3 more
Multiple sequence alignment (MSA) inference is a central task in molecular evolution and comparative genomics, and the reliability of downstream analyses, including phylogenetic inference, depends critically on alignment quality. Despite this importance, most widely used MSA methods optimize the sum-of-pairs (SoP)…
Paul A Gagniuc, Elvira Gagniuc
Sequence alignment provides a formal framework for comparison of biological sequences through score maximization over matches, mismatches, and insertion-deletion events. Classical formulations distinguish between global alignment, which enforces end-to-end correspondence through fixed boundary conditions, and local…
Emily G. Light, Morgan E. Prior, Noah M. Daniels, Najib Ishaq
Motivation: The multiple sequence alignment (MSA) problem has been extensively studied, with numerous approaches developed over recent years. With the rapid growth of sequence data, there is an increasing need for fast and accurate MSA tools that scale effectively to large datasets. Building on our previous work on…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
Background The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated…
Yoshiki Kanazawa, Naphan Benchasattabuse, Michal Hajdušek, Rodney Van Meter
We formulate a structure-informed multiple sequence alignment problem, denoted MSA-S. The model abstracts biological sequences as strings and structural information as designated position-pairs. It augments a fixed pairwise string score, defined by a fixed non-gap symbol-pair scoring rule and fixed affine gap…
Daniel Yang, Thaxter Shaw, TJ Tsai
This article investigates several parallelizable alternatives to DTW for estimating the alignment between two long sequences. Whereas most previous work has focused on reducing the total computation and/or memory costs of DTW, our focus is instead on reducing wall clock time by utilizing common hardware like GPUs that…
Lasse Reifenrath, Michel van Kempen, Gyuri Kim, Soo Hyun Kim + 4 more
The ubiquitous availability of protein structures permits replacing sequence alignment with more accurate and sensitive structure alignment algorithms. LoL-align maximizes a local log-odds score for proteins to be homologous, given their intra-protein C_α_ – C_α_ distances. LoL-align is markedly more sensitive in…
Muhammad Shoaib, Waqas Ali
Dynamic programming (DP) yields exact quadratic-time (O(NM)) pairwise sequence alignments. Static banding heuristics (O(NW)) fail catastrophically on low-identity (< 30%), asymmetric insertions/deletions (indels), or extreme length ratios, dropping core-block Sum-of-Pairs (SP) score recovery to 20%– 50%. Conversely…
Simeng Zhang, Xinying Liu, Jun Lou, Mudi Jiang + 2 more
—The rapid development of high-throughput sequencing technologies has led to an explosive increase in biological sequence data, making sequence clustering a fundamental task in large-scale bioinformatics analyses. Unlike traditional clustering problems, biological sequence clustering faces unique challenges due to the…
Boryeu Mao
Sequence matching algorithms such as BLAST and FASTA have been widely used in searching for evolutionary origin and biological functions of newly discovered nucleic acid and protein sequences. As parts of these search tools, alignment scores and E values are useful indicators of the quality of search results (and the…
Jingqing Hu, Qian Qin, Heng Li, Ying Zhou + 1 more
A pair of template alleles for each targeted gene are used as references to construct consensus allele sequences from informative reads, based on the assumption that the target sample has exact two copies of each of the six genes. Template allele pair is a combination of reference alleles of the highest agreement with…
Authors not listed
Computational modeling of enzymes provides molecular-level insight into catalysis, but the preparation of quantum mechanical (QM) calculations starting from experimental structures is a significant bottleneck for high-throughput studies. Automated tools developed to accelerate this process may fail to generalize across…
Authors not listed
Monoterpene synthases (mTSs) are a large family of enzymes, which have promising industrial applications, yet remain difficult to engineer due to complex and poorly understood sequence-function relationships. Here, we present a structure-based machine learning (ML) framework that accurately predicts whether a mTS…
Authors not listed
A framework for catalysis based on categorical aperture selection rather than temporal acceleration is presented. Traditional catalysis theory describes catalysts as agents that accelerate reactions by lowering activation energies, implicitly treating time as the fundamental variable and reaction rate enhancement as…
Authors not listed
RNA molecules fold into complex three-dimensional structures that determine their function. A wide range of mathematical frameworks, such as chord diagrams, fatgraphs, and context-free grammars, have been used to represent these structures; however, these models have largely been developed from mathematical motivations…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Authors not listed
The development of biomaterials that mimic the native extracellular matrix (ECM) of native tissue represents an exciting frontier for tissue engineering and regenerative medicine. Injectable nano-fiber hydrogels made of short, self-assembling peptides offer a promising platform for the delivery and directed…
Authors not listed
Accurately measuring compound binding affinities is key to driving the pharmaceutical development process. Rigorous physics-based in silico approaches, particularly alchemical free energy methods, have become a gold standard tool for estimating compound affinity changes. Here we present the results of a large-scale…