14 papers · ranked by Valyu relevance
Daniel Liu, Martin Steinegger
The Smith-Waterman-Gotoh alignment algorithm is the most popular method for comparing biological sequences. Recently, Single Instruction Multiple Data methods have been used to speed up alignment. However, these algorithms have limitations like being optimized for specific scoring schemes, cannot handle large gaps, or…
Peiyu Zong, Wenpeng Deng, Jian Liu, Jue Ruan
The rapid advancements in sequencing length necessitate the adoption of increasingly efficient sequence alignment algorithms. The Needleman-Wunsch method introduces the foundational dynamic programming (DP) matrix calculation for global alignment, which evaluates the overall alignment of sequences. However, this method…
Pesho Ivanov, Benjamin Bichsel, Martin Vechev
We present a novel A^⋆^ seed heuristic that enables fast and optimal sequence-to-graph alignment, guaranteed to minimize the edit distance of the alignment assuming non-negative edit costs. We phrase optimal alignment as a shortest path problem and solve it by instantiating the A^⋆^ algorithm with our seed heuristic.…
Haojing Shao, Jue Ruan
Increasing the accuracy of the nucleotide sequence alignment is an essential issue in genomics research. Although classic dynamic-programming algorithms (e.g., Smith-Waterman and Needleman–Wunsch) guarantee to produce the optimal result, their time complexity hinders the application of large-scale sequence alignment.…
Jordan M. Eizenga, Benedict Paten
Modern genomic sequencing data is trending toward longer sequences with higher accuracy. Many analyses using these data will center on alignments, but classical exact alignment algorithms are infeasible for long sequences. The recently proposed WFA algorithm demonstrated how to perform exact alignment for long, similar…
Santhosh Sankar, Naren Chandran Sakthivel, Nagasuma Chandra
Protein function is a direct consequence of its sequence, structure and the arrangement at the binding site. Bioinformatics using sequence analysis is typically used to gain a first insight into protein function. Protein structures, on the other hand, provide a higher resolution platform into understanding functions.…
Julia Malec, Karina Rusen, G. Brian Golding, Lucian Ilie
Protein sequence alignment is one of the most fundamental procedures in bioinformatics. Due to its many downstream applications, improvements to this procedure are of great importance. We consider two revolutionary concepts that emerged recently as candidates for improving the state-of-the-art alignment methods…
Zhengyang Guo, Yang Wang, Guangshuo Ou
Protein structure comparison is pivotal for deriving homological relationships, elucidating protein functions, and understanding evolutionary developments. The burgeoning field of in-silico protein structure prediction now yields billions of models with near-experimental accuracy, necessitating sophisticated tools for…
Ragnar Groot Koerkamp, Igor Martayan
Because of the rapidly-growing amount of sequencing data, computing sketches of large textual datasets has become an essential preprocessing task. These sketches are typically much smaller than the input sequences, but preserve sufficient information for downstream analysis. Minimizers are an especially popular…
Saba Zerefa, Jesse Cool, Pramesh Singh, Samantha Petti
Recent advancements in protein structure prediction methods have vastly increased the size of databases of protein structures, necessitating fast methods for protein structure comparison. Search methods that find structurally similar proteins can be applied to find remote homologs, study the functional relationships…
Rahul Varki, Christina Boucher
Relative Lempel–Ziv (RLZ) is an effective compression method for large, repetitive collections; however, the fundamental primitives required to elevate it from a passive archival format to a tractable representation for compressed construction have yet to be fully established. In this paper, we introduce an algorithmic…
Maria Evangelia Vlachou, Elizabeth Thomas, Jean Blouin
In this paper, we address the problem of quantifying similarity between planar 2D shapes, which is relevant to studies of internal representations in cognitive, developmental, and neurological research. We designed a set of test shapes arranged along a visually defined perceptual similarity gradient and used them to…
Daniel J. van Zyl, Marcel Dunaiski, Houriiyah Tegally, Cheryl Baxter + 2 more
The rapid increase in nucleotide sequence data generated by next-generation sequencing (NGS) technologies demands efficient computational tools for sequence comparison. Alignment-based methods, such as BLAST, are increasingly overwhelmed by the scale of contemporary datasets due to their high computational demands for…
Andreas Grigorjew, Artur Gynter, Fernando Dias, Benjamin Buchfink + 2 more
Sequence alignments have become the foundation of life science research by unlocking biological mechanisms through protein comparisons. Despite its methodological success, most algorithmic innovation in the past decades focused on the optimal alignment problem, while often ignoring information derived from suboptimal…