18 papers · ranked by Valyu relevance
Patrick Diehl, Noujoud Nader, Deepti Gupta
Parallel programming remains one of the most challenging aspects of High-Performance Computing (HPC), requiring deep knowledge of synchronization, communication, and memory models. While modern C++ standards and frameworks like OpenMP and MPI have simplified parallelism, mastering these paradigms is still complex.…
Shiting Long, Gustavo Ramirez-Hidalgo, Andreas Frommer, Dirk Pleiter
Gauss-Seidel is a well-established iterative method for the solution of linear systems, and multicoloring has been widely used to increase parallelism in iterative solution techniques. Implementing multi-color Gauss-Seidel with conventional divide-and-conquer parallelization strategies, however, may be inefficient due…
Haymo Kutschbach
This work introduces a self-optimizing virtual processor (VP) for numerical array programs that shifts parallelization from a manual developer task to a cooperative, agent-like runtime mechanism. Instead of relying on centralized task-graph scheduling, static compiler optimization, or explicitly annotated parallel…
William B. Langdon
—We extend recent 256 SSE vector work to 512 AVX giving a four fold speedup. We use MAGPIE (Machine Automated General Performance Improvement via Evolution of software) to speedup a C++ linear genetic programming interpreter. Local search is provided with three alternative hand optimised codes, revision history and the…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Villalobos, Johansell, Ruzicka, Josef + 2 more
—Scientific computing in the exascale era demands increased computational power to solve complex problems across various domains. With the rise of heterogeneous computing architectures the need for vendor-agnostic, performance portability frameworks has been highlighted. Libraries like Kokkos have become essential for…
Ran Ginosar
I have greatly enjoyed spending many years in studying parallel computing. My journey goes thorough MP-C, PLURAL, Async Plural, HAL, RC64 and more. As a PhD student at Princeton I studied a combination of shared memory and message passing, motivated by algorithms and the ease of programming. While at the Technion, a…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Linus Bantel, Anna-Lena Roth, Jonas Posner, Dirk Pflüger
Julia is increasingly used in hpc as a single-language alternative to combining high-level scripting with low-level systems languages, but achieving scalable performance still requires expertise in parallel programming. llms are increasingly used for code generation and are advancing rapidly with each new version. Yet…
Mateusz Gruzewski, Marek Palkowski, Ramon Antonio Rodriges Zalipynis
In this article, we present an efficient and concise OpenMP implementation of the Nussinov RNA folding algorithm, a well-known representative of non-serial polyadic dynamic programming (NPDP). Our goal is to develop an optimized implementation that can serve as a template for related dynamic programming applications.…
Christoph Stelz, Lukas Hübner, Alexandros Stamatakis, Russell Schwartz
Given an input array of elements $E=[e_{0},e_{1},…,e_{n-1}]$, distributed across p processing elements (PEs; e.g. processes or threads), we desire to compute $r=e_{0}⊕e_{1}⊕…⊕e_{n-1}$, where $⊕$ denotes a binary, associative operation (e.g. summation or multiplication). In a distributed reduction, we return the result…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
Marek Wiewiórka, Pavel Khamutou, Marek Zbysiński, Tomasz Gambin + 1 more
Operations on genomic intervals (ranges), such as finding intervals that overlap with or are closest to a given interval, are fundamental to many bioinformatics analyses, including the process of annotating genetic variants (). Dedicated efficient data structures and/or accompanying command-line interface tools, such…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
Background The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated…
Anders Pitman, Cathy Yang, Yi Qiao
Next-generation sequencing now produces whole-genome data in hours, but downstream variant calling remains a multi-hour to multi-day bottleneck that excludes genomic analysis from time-critical clinical settings. GPU acceleration offers a natural path forward — variant calling is inherently parallelizable across…
Jian Liu, Jianwei Zhang
Current de novo genome assembly tools often demand substantial memory resources, and their execution typically relies on high-performance computing (HPC) clusters. This dependency limits their use in resource-constrained settings. Furthermore, mainstream third-generation sequencing assembly and alignment tools usually…