14 papers · ranked by Valyu relevance
G. Kandemir, D. H. Duncan, D. van Moorselaar, J. Theeuwes
For almost half a century, target-distractor similarity has been known to induce different visual search modes. When a target is highly salient, it can pop out, suggesting parallel processing of all items irrespective of set size. By contrast, high similarity among items requires item-by-item comparison with an…
Tim Anderson, Travis J Wheeler
Pattern matching is a key step in a variety of biological sequence analysis pipelines. The FM-index is a compressed data structure for pattern matching, with search run time that is independent of the length of the database text. We present AvxWindowedFMindex (AWFM-index), an open-source, thread-parallel FM-index…
Sumesh Kumar, Joseph Zambreno, Ashfaq Khokhar, Shoaib Akram + 1 more
Improving the speed and efficiency of database search algorithms that deduce peptides from mass spectrometry (MS) data has been an active area of research for more than three decades. The significance of the need for faster database search methods has rapidly increased due to the growing interest in studying non-model…
Bertil Schmidt, Felix Kallenborn, Alejandro Chacon, Christian Hundt
The maximal sensitivity for local pairwise alignment makes the Smith-Waterman algorithm a popular choice for protein sequence database search. However, its quadratic time complexity makes it compute-intensive. Unfortunately, current state-of-the-art software tools are not able to leverage the massively parallel…
Wilfried Agbeto, Camille Coti, Vladimir Reinharz
Advances in graph algorithmics have allowed in-depth study of many natural objects from molecular biology or chemistry to social networks. Particularly in molecular biology and cheminformatics, understanding complex structures by identifying conserved sub-structures is a key milestone towards the artificial design of…
Evelin Aasna, Simon Gene Gottlieb, Marcel Ehrhardt, Knut Reinert
Searching large genomic data sets for local alignments poses a computational challenge. A particular obstacle is the handling of repetitive sequences that appear in various contexts and incur a high runtime cost. For practical homology search, it is important to develop a specific but sensitive filter. Good filters…
Stuart Byma, Akash Dhasade, Adrian Altenhoff, Christophe Dessimoz + 1 more
This paper presents a new, parallel implementation of clustering and demonstrates its utility in greatly speeding up the process of identifying homologous proteins. Clustering is a technique to reduce the number of comparison needed to find similar pairs in a set of n elements such as protein sequences. Precise…
Peiyu Zong, Wenpeng Deng, Jian Liu, Jue Ruan
The rapid advancements in sequencing length necessitate the adoption of increasingly efficient sequence alignment algorithms. The Needleman-Wunsch method introduces the foundational dynamic programming (DP) matrix calculation for global alignment, which evaluates the overall alignment of sequences. However, this method…
Kecong Tang, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang
Durbin’s positional Burrows-Wheeler transform (PBWT) enables algorithms with the optimal time complexity of O(MN) for reporting all vs all haplotype matches in a population panel with M haplotypes and N variant sites. However, even this efficiency may still be too slow when the number of haplotypes reaches millions. To…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated quadratic…
Richard Wilton, Tamas Budavari, Ben Langmead, Sarah Wheelan + 2 more
In computing pairwise alignments of biological sequences, software implementations employ a variety of heuristics that decrease the computational effort involved in computing potential alignments. A key element in achieving high processing throughput is to identify and prioritize potential alignments where high-scoring…
Seth Stadick
Filtering records using command line tools is a staple of Bioin-formatics. In analysis pipelines and in day-to-day research tools such as awk, grep, and cut are the workhorses of much of our data crunching. To date, there is no command line utility for performing index-free alignment-based filtering of records. Ish is…
Christoph Stelz, Lukas Hübner, Alexandros Stamatakis
Phylogenetic trees describe the evolutionary history among biological species based on their genomic data. Maximum Likelihood (ML) based phylogenetic inference tools search for the tree and evolutionary model that best explain the observed genomic data. Given the independence of likelihood score calculations between…
Costanza Pascal, Herzeel Charlotte, Verachtert Wilfried
elPrep is an established multi-threaded framework for preparing SAM and BAM files in sequencing pipelines. To achieve good performance, its software architecture makes only a single pass through a SAM/BAM file for multiple preparation steps, and keeps sequencing data as much as possible in main memory. Similar to other…