13 papers · ranked by Valyu relevance
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Seth Stadick
Filtering records using command line tools is a staple of Bioin-formatics. In analysis pipelines and in day-to-day research tools such as awk, grep, and cut are the workhorses of much of our data crunching. To date, there is no command line utility for performing index-free alignment-based filtering of records. Ish is…
Zeyu Xia, Canqun Yang, Chenchen Peng, Yifei Guo + 3 more
'Tao Tang' 'Yingbo Cui'] Background The advent of Single Molecule Real-Time (SMRT) sequencing has overcome many limitations of second-generation sequencing, such as limited read lengths, PCR amplification biases. However, longer reads increase data volume exponentially and high error rates make many existing alignment…
Paul Cardosi, Bérenger Bramas, Bilal Alatas
Parallelization is needed everywhere, from laptops and mobile phones to supercomputers. Among parallel programming models, task-based programming has demonstrated a powerful potential and is widely used in high-performance scientific computing. Not only does it allow efficient parallelization across distributed…
Marissa E. Powers, Keith Mannthey, Priyanka Sebastian, Snehal Adsule + 6 more
Next Generation Sequencing (NGS) workloads largely consist of pipelines of tasks with heterogeneous compute, memory, and storage requirements. Identifying the optimal system configuration has historically required expertise in both system architecture and bioinformatics. This paper outlines infrastructure…
Peiyu Zong, Wenpeng Deng, Jian Liu, Jue Ruan
The rapid advancements in sequencing length necessitate the adoption of increasingly efficient sequence alignment algorithms. The Needleman-Wunsch method introduces the foundational dynamic programming (DP) matrix calculation for global alignment, which evaluates the overall alignment of sequences. However, this method…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated quadratic…
Sirilak Ketchaya, Apisit Rattanatranurak
Quicksort is an important algorithm that uses the divide and conquer concept, and it can be run to solve any problem. The performance of the algorithm can be improved by implementing this algorithm in parallel. In this paper, the parallel sorting algorithm named the Multi-Deque Partition Dual-Deque Merge Sorting…
Ashish Chapagain, Dima Abuoliem, In Ho Cho, Tongbiao Wang
Multifunctional nanosurfaces receive growing attention due to their versatile properties. Capillary force lithography (CFL) has emerged as a simple and economical method for fabricating these surfaces. In recent works, the authors proposed to leverage the evolution strategies (ES) to modify nanosurface characteristics…
Vladislav Skorpil, Vaclav Oujezsky, Arcangelo Castiglione, Gianni D’Angelo
'Gianni D’Angelo'] This paper presents an implementation of the parallelization of genetic algorithms. Three models of parallelized genetic algorithms are presented, namely the Master-Slave genetic algorithm, the Coarse-Grained genetic algorithm, and the Fine-Grained genetic algorithm. Furthermore, these models are…
Kecong Tang, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang
Durbin’s positional Burrows-Wheeler transform (PBWT) enables algorithms with the optimal time complexity of O(MN) for reporting all vs all haplotype matches in a population panel with M haplotypes and N variant sites. However, even this efficiency may still be too slow when the number of haplotypes reaches millions. To…
Fayez AlFayez, Muhammad Aleem
This work investigates minimizing the makespan of multiple servers in the case of identical parallel processors. In the case of executing multiple tasks through several servers and each server has a fixed number of processors. The processors are generally composed of two processors (core duo) or four processors (quad).…
Patrick McKeever, Varun Mittal, Bryce Fukuda, Ka Yee Yeung + 1 more
The exponential growth of omics data requires novel strategies for storage, transfer, and processing of said data. We present a scheduler based on the Temporal.io workflow framework which enables two key optimizations of bioinformatics workflows. Firstly, we enable users to transparently map workflow steps to diverse…