Search · four archives
Search · four archives
24 papers · ranked by Valyu relevance
Brandon Legried, Sébastien Roch
Ancestral sequence reconstruction is a key task in computational biology. It consists in inferring a molecular sequence at an ancestral species of a known phylogeny, given descendant sequences at the tip of the tree. In addition to its many biological applications, it has played a key role in elucidating the…
Liang Hong, Siqi Sun, Liangzhen Zheng, Qingxiong Tan + 1 more
Evolutionarily related sequences provide information for the protein structure and function. Multiple sequence alignment, which includes homolog searching from large databases and sequence alignment, is efficient to dig out the information and assist protein structure and function prediction, whose efficiency has been…
Emily G. Light, Morgan E. Prior, Noah M. Daniels, Najib Ishaq
Motivation: The multiple sequence alignment (MSA) problem has been extensively studied, with numerous approaches developed over recent years. With the rapid growth of sequence data, there is an increasing need for fast and accurate MSA tools that scale effectively to large datasets. Building on our previous work on…
Jiannan Chao, Furong Tang, Lei Xu, Lukasz Kurgan
The continuous development of sequencing technologies has enabled researchers to obtain large amounts of biological sequence data, and this has resulted in increasing demands for software that can perform sequence alignment fast and accurately. A number of algorithms and tools for sequence alignment have been designed…
Claire D. McWhite, Mona Singh
Multiple sequence alignment is a critical step in the study of protein sequence and function. Typically, multiple sequence alignment algorithms progressively align pairs of sequences and combine these alignments with the aid of a guide tree. These alignment algorithms use scoring systems based on substitution matrices…
Petar Arsic, Christoph Mayer
We report a convolutional transformer neural network that is capable of aligning multiple nucleotide sequences. The neural network is based on the U-Net commonly used in image segmentation which we employ to transform unaligned sequences to aligned sequences. For alignment scenarios our Ali-U-Net neural network has…
Yanming Wei, Zhaoyang Huang, Pinglu Zhang, Yizheng Wang + 3 more
Multiple sequence alignment (MSA) is a fundamental problem in bioinformatics. The quality of sequence alignment significantly impacts biological sequence analysis, especially that in next-generation sequencing . MSA results are widely used in various applications, including de novo genome assembly , detection of…
Nimrod Serok, Ksenia Polonsky, Haim Ashkenazy, Itay Mayrose + 3 more
Multiple sequence alignment (MSA) inference is a central task in molecular evolution and comparative genomics, and the reliability of downstream analyses, including phylogenetic inference, depends critically on alignment quality. Despite this importance, most widely used MSA methods optimize the sum-of-pairs (SoP)…
Yoshiki Kanazawa, Naphan Benchasattabuse, Michal Hajdušek, Rodney Van Meter
We formulate a structure-informed multiple sequence alignment problem, denoted MSA-S. The model abstracts biological sequences as strings and structural information as designated position-pairs. It augments a fixed pairwise string score, defined by a fixed non-gap symbol-pair scoring rule and fixed affine gap…
Bryce Kille, Advait Balaji, Fritz J. Sedlazeck, Michael Nute + 1 more
'Todd J. Treangen'] With the arrival of telomere-to-telomere (T2T) assemblies of the human genome comes the computational challenge of efficiently and accurately constructing multiple genome alignments at an unprecedented scale. By identifying nucleotides across genomes which share a common ancestor, multiple genome…
Dimitrii O. Kostenko, Eugene V. Korotkov, Cristoforo Comi, Benoit Gauthier + 2 more
'Benoit Gauthier' 'Dimitrios H. Roukos' 'Alfredo Fusco'] The aim of this work was to compare the multiple alignment methods MAHDS, T-Coffee, MUSCLE, Clustal Omega, Kalign, MAFFT, and PRANK in their ability to align highly divergent amino acid sequences. To accomplish this, we created test amino acid sequences with an…
Luisa Santus, Jose Espinosa-Carrasco, Leon Rauschning, Júlia Mir-Pedrol + 12 more
The computational complexity of many key bioinformatics problems has resulted in numerous alternative heuristic solutions, where no single approach consistently outperforms all others. This creates difficulties for users trying to identify the most suitable tool for their dataset and for developers managing and…
Luisa Santus, Jose Espinosa-Carrasco, Leon Rauschning, Júlia Mir-Pedrol + 12 more
'Júlia Mir-Pedrol' 'Igor Trujnara' 'Alessio Vignoli' 'Leila Mansouri' 'Athanasios Baltzis' 'Evan W Floden' 'Paolo Di\xa0Tommaso' 'Edgar Garriga' 'Adam Gudyś' 'Sebastian Deorowicz' 'Cameron Gilchrist' 'Martin Steinegger' 'Cedric Notredame'] Title: Abstract The computational complexity of many key bioinformatics problems…
Suchindra, P. Nagaraj
DNA sequence alignment is important today as it is usually the first step in finding gene mutation, evolutionary similarities, protein structure, drug development and cancer treatment. Covid-19 is one recent example. There are many sequencing algorithms developed over the past decades but the sequence alignment using…
Michail Patsakis, Kimonas Provatas, Fotis A. Baltoumas, Nikol Chantzi + 3 more
'Nikol Chantzi' 'Ioannis Mouratidis' 'Georgios A. Pavlopoulos' 'Ilias Georgakopoulos-Soares'] Motivation: Genome and Proteome Alignments, represented by the Multiple Alignment File (MAF) format, have become a standard approach in the field of comparative genomics and proteomics. However, current approaches lack a…
Hao Xuan, Hongyang Sun, Xiangtao Liu, Hanyuan Zhang + 2 more
Sequence alignment underpins nearly every facet of modern genomics, from genetic testing and cancer profiling to functional genome annotation. Yet, despite decades of algorithmic innovation, most existing aligners remain narrowly optimized for specific tasks, fragmenting analytical workflows and limiting…
Xinwei Zhao, Eberhard Korsching, Philip Hublitz
DNA and RNA nucleotide sequences are ubiquitous in all biological cells, serving as both a comprehensive library of capabilities for the cells and as an impressive regulatory system to control cellular function. The multi-alignment framework (MAF) provided in this study offers a user-friendly platform for sequence…
Xinwei Zhao, Eberhard Korsching
DNA and RNA nucleotide sequences are ubiquitous in all biological cells, serving as both a comprehensive library of capabilities for the cells and as an impressive regulatory system to control cellular function. The multi-alignment framework (MAF) provided in this study offers a user-friendly platform for sequence…
Brendan Furneaux, Sten Anslan, Panu Somervuo, Jenni Hultman + 3 more
and a bioinformatics workflow for metabarcoding data Authors: ['Brendan Furneaux' 'Sten Anslan' 'Panu Somervuo' 'Jenni Hultman' 'Nerea Abrego' 'Tomas Roslin' 'Otso Ovaskainen'] To turn environmentally derived metabarcoding data into community matrices for ecological analysis, sequences must first be clustered into…
Authors not listed
The Protein Data Bank (PDB) is one of the richest open‑source repositories in biology, housing over 277,000 macromolecular structural models alongside much of the experimental data that underpins these models. By systematically collecting, validating, and indexing these models, the PDB has accelerated structural…
Pavel Kohout, Michal Vasina, Marika Majerova, Veronika Novakova + 5 more
Enzymes play a crucial role in sustainable industrial applications, with their optimization posing a formidable challenge due to the intricate interplay among residues. Computational methodologies predominantly rely on evolutionary insights, leveraging homologous sequences to pinpoint conserved and functionally…
Authors not listed
Proteochemometric models (PCM) are used in computational drug discovery to leverage both protein and ligand representations for bioactivity prediction. While machine learning (ML) and deep learning (DL) have come to dominate PCMs, often serving as scoring functions, rigorous evaluation standards have not always been…
Joseph Redshaw, Darren Ting, Alex Brown, Jonathan Hirst + 1 more
Antimicrobial peptides (AMPs) represent a potential solution to the growing problem of antimicrobial resistance, yet their identification through wet-lab experiments is a costly and timeconsuming process. Accurate computational predictions would allow rapid in silico screening of candidate AMPs, thereby accelerating…
Babu Bassa
In this communication the author describes a software tool named "ChameleonSort". The software program, developed by the present author is useful in the sorting of biological sequence variants like those accumulating mutations while diverging from the common ancestors. Examples include viral protein variants, protein…