25 papers · ranked by Valyu relevance
Andre Müller, Alexander Wichmann, Felix Kallenborn, Andreas Hildebrandt + 2 more
Each read is aligned against all candidate regions obtained in the query step using semi-global alignment methods provided by Edlib (). Using the alignment information we then construct an MSA of a candidate region and all reads mapped to that region in the reference. This enables us to remove certain false positive…
Albert Jiménez-Blanco, Lorién López-Villellas, Juan Carlos Moure, Miquel Moreto + 1 more
Sequence-to-graph alignment is a central problem in bioinformatics, with applications in multiple sequence alignment (MSA) and pangenome analysis, among others. However, current algorithms for optimal affine-gap alignment impose high memory and computational requirements, limiting their scalability to aligning long…
Angelica Lindlöf
Next-generation sequencing (NGS) is a technology that enables rapid and high-throughput sequencing of entire genomes, transcriptomes or specific DNA/RNA populations. RNA-Seq is an NGS-based method that specifically targets the transcriptome and can be applied to bulk tissue or single cells. NGS produces large volumes…
Mohammed Alser, Julien Eudine, Onur Mutlu
Searching for similar genomic sequences is an essential and fundamental step in biomedical research and an overwhelming majority of genomic analyses. State-of-the-art computational methods performing such comparisons fail to cope with the exponential growth of genomic sequencing data. We introduce the concept of…
Kristoffer Sahlin
Read alignment is often the computational bottleneck in analyses. Recently, several advances have been made on seeding methods for fast sequence comparison. We combine two such methods, syncmers and strobemers, in a novel seeding approach for constructing dynamic-sized fuzzy seeds and implement the method in a…
Tommi Mäklin, Jarno N. Alanko, Elena Biagi, Simon J. Puglisi
Finding high-quality local alignments between a query sequence and sequences contained in a large genomic database is a fundamental problem in computational genomics, at the core of thousands of biological analysis pipelines. Here, we describe a novel algorithm for approximate local alignment search based on the…
Alisa Prusokiene, Neil Boonham, Adrian Fox, Thomas P. Howard + 1 more
'Tarunendu Mapder'] Current tools for estimating the substitution distance between two related sequences struggle to remain accurate at a high divergence. Difficulties at distant homologies, such as false seeding and over-alignment, create a high barrier for the development of a stable estimator. This is especially…
Hao Xuan, Hongyang Sun, Xiangtao Liu, Hanyuan Zhang + 2 more
Sequence alignment underpins nearly every facet of modern genomics, from genetic testing and cancer profiling to functional genome annotation. Yet, despite decades of algorithmic innovation, most existing aligners remain narrowly optimized for specific tasks, fragmenting analytical workflows and limiting…
Robert C. Edgar
Recent breakthroughs in protein fold prediction from amino acid sequences have unleashed a deluge of new structures, raising new opportunities for expanding insights into the universe of proteins and pursuing practical applications in bio-engineering and therapeutics while also presenting new challenges to protein…
Przemysław Stawczyk, Robert Nowak
Reducing the cost of sequencing genomes provided by next-generation sequencing technologies has greatly increased the number of genomic projects. As a result, there is a growing need for better assembly and assembly validation methods. One promising idea is to use heterogeneous data in assembly projects. Optical…
Sung Jong Lee, Keehyoung Joo, Sangjin Sim, Juyong Lee + 3 more
'Jooyoung Lee' 'Michael Assfalg'] Sequence-structure alignment for protein sequences is an important task for the template-based modeling of 3D structures of proteins. Building a reliable sequence-structure alignment is a challenging problem, especially for remote homologue target proteins. We built a method of…
Kejue Jia, Mesih Kilinc, Robert L. Jernigan
Understanding protein sequences and how they relate to the functions of proteins is extremely important. One of the most basic operations in bioinformatics is sequence alignment and usually the first things learned from these are which positions are the most conserved and often these are critical parts of the…
Mingeun Ji, Yejin Kan, Dongyeon Kim, Jaehee Jung + 2 more
'Frank M. You'] Advances in the next-generation sequencing technology have led to a dramatic decrease in read-generation cost and an increase in read output. Reconstruction of short DNA sequence reads generated by next-generation sequencing requires a read alignment method that reconstructs a reference genome. In…
Christos Argyropoulos
component based applications using Object Orientation, PDL, Alien, FFI, Inline and OpenMP Authors: ['Christos Argyropoulos'] Component-Based Software Engineering (CBSE) is a methodology that assembles pre-existing, reusable software components into new applications, which is particularly relevant for fast moving…
Suchindra, P. Nagaraj
DNA sequence alignment is important today as it is usually the first step in finding gene mutation, evolutionary similarities, protein structure, drug development and cancer treatment. Covid-19 is one recent example. There are many sequencing algorithms developed over the past decades but the sequence alignment using…
Alsamman M. Alsamman, Achraf El Allali, Morad M. Mokhtar, Khaled Al-Sham’aa + 4 more
'Khaled Al-Sham’aa' 'Ahmed E. Nassar' 'Khaled H. Mousa' 'Zakaria Kehel' 'Nikolas Pontikos'] Multiple sequence alignment (MSA) is essential for understanding genetic variations controlling phenotypic traits in all living organisms. The post-analysis of MSA results is a difficult step for researchers who do not have…
Olga Ivanova, Jose Gavaldá-García, Dea Gogishvili, Isabel Houtkamp + 3 more
'Robbin Bouwmeester' 'K. Anton Feenstra' 'Sanne Abeln'] | 3 | Structure Alignment | | 1 | | --- | --- | --- | --- | | | Olga Ivanova | Jose Gavald´a-Garc´ıa Dea Gogishvili | | | | Isabel Houtkamp | Robbin Bouwmeester | | | | K. Anton Feenstra | Sanne Abeln | | | | 1 Comparing protein structures | | 4 | | | | 1.1…
Yoshiki Kanazawa, Naphan Benchasattabuse, Michal Hajdušek, Rodney Van Meter
We formulate a structure-informed multiple sequence alignment problem, denoted MSA-S. The model abstracts biological sequences as strings and structural information as designated position-pairs. It augments a fixed pairwise string score, defined by a fixed non-gap symbol-pair scoring rule and fixed affine gap…
Manal Helal, Fanrong Kong, Sharon C.‐A. Chen, Fei Zhou + 3 more
'Dominic E. Dwyer' 'John Potter' 'Vitali Sintchenko'] Background: Comparative genomics has put additional demands on the assessment of similarity between sequences and their clustering as means for classification. However, defining the optimal number of clusters, cluster density and boundaries for sets of potentially…
Chengze Shen, Baqiao Liu, Kelly P Williams, Tandy Warnow
Adding sequences into an existing (possibly user-provided) alignment has multiple applications, including updating a large alignment with new data, adding sequences into a constraint alignment constructed using biological knowledge, or computing alignments in the presence of sequence length heterogeneity. Although this…
Xinwei Zhao, Eberhard Korsching
DNA and RNA nucleotide sequences are ubiquitous in all biological cells, serving as both a comprehensive library of capabilities for the cells and as an impressive regulatory system to control cellular function. The multi-alignment framework (MAF) provided in this study offers a user-friendly platform for sequence…
Edgar López-López, Oscar Robles, Fabien Plisson, José L. Medina-Franco
Peptides are a re-emerged strategy to fight a plethora of diseases and their utility has been expanded to new areas. Now sequence-based peptide design opens up new possibilities to develop peptidic molecular entities. However, its methodological limitations (e.g., its inefficiency in designing large peptides and that…
Joshua Meyers, Nathan Brown
De novo molecular design is effective for optimizing compounds with desirable properties in drug discovery. In the realm of chemoinformatics, population-based evolutionary approaches have been known for over a decade including those that use retrosynthetically inspired fragment-based workflows to to balance the…
Simeng Zhang, Xinying Liu, Jun Lou, Mudi Jiang + 2 more
—The rapid development of high-throughput sequencing technologies has led to an explosive increase in biological sequence data, making sequence clustering a fundamental task in large-scale bioinformatics analyses. Unlike traditional clustering problems, biological sequence clustering faces unique challenges due to the…
Joseph Redshaw, Darren Ting, Alex Brown, Jonathan Hirst + 1 more
Antimicrobial peptides (AMPs) represent a potential solution to the growing problem of antimicrobial resistance, yet their identification through wet-lab experiments is a costly and timeconsuming process. Accurate computational predictions would allow rapid in silico screening of candidate AMPs, thereby accelerating…