11 papers · ranked by Valyu relevance
de Ronde, Folkert, Knapen, Alexander + 4 more
This paper introduces a novel approach to enhance the execution of quantum algorithms on distributed quantum systems. The proposed method involves the development of a hardware design that supports parallel instruction execution and a compiler that modifies the order of instructions to increase parallelism…
Haymo Kutschbach
This work introduces a self-optimizing virtual processor (VP) for numerical array programs that shifts parallelization from a manual developer task to a cooperative, agent-like runtime mechanism. Instead of relying on centralized task-graph scheduling, static compiler optimization, or explicitly annotated parallel…
Aleix Roca, Vicenç Beltran
The convergence of high-performance computing (HPC) and artificial intelligence (AI) is driving the emergence of increasingly complex parallel applications and workloads. These workloads often combine multiple parallel runtimes within the same application or across co-located jobs, creating scheduling demands that…
Kecong Tang, Ardalan Naseri, Degui Zhi, Shaojie Zhang + 1 more
To support memory-efficient execution without compromising the all-vs.-all nature of IBD detection, RaPID2 adopts a partitioning strategy that operates on haplotype pairs () rather than subpanels. All pairs are normalized such that $A<B$ to ensure consistent key assignment, and each pair is deterministically assigned…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
Harry Fitchett, Charles M. Fox
While transistor density is still increasing, clock speeds are not, motivating the search for new parallel architectures. One approach is to completely abandon the concept of CPU – and thus serial imperative programming – and instead to specify and execute tasks in parallel, compiling from programming languages to data…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Alexander Strack, Alexander Van Craen, Dirk Pflüger
Fork-join parallelism, popularized by OpenMP, remains the dominant model for shared-memory parallel programming, but its implicit synchronization barriers can penalize algorithms with inhomogeneous workloads. Asynchronous many-task (AMT) runtimes sidestep these barriers by expressing work as a dependency graph of…
John Kruper, Ariel Rokem
Tractography based on diffusion-weighted MRI (dMRI) is the predominant in vivo method for mapping the brain’s white matter. However, it is also one of the most computationally demanding steps in neuroimaging data analysis-requiring the generation and filtering of millions of streamlines per subject. Over the past…
Mateusz Gruzewski, Marek Palkowski, Ramon Antonio Rodriges Zalipynis
In this article, we present an efficient and concise OpenMP implementation of the Nussinov RNA folding algorithm, a well-known representative of non-serial polyadic dynamic programming (NPDP). Our goal is to develop an optimized implementation that can serve as a template for related dynamic programming applications.…
Ismail Melik Turker, Isa Yildirim
3.3.1#### Experimental setup All experiments were conducted on a laptop running Ubuntu 24.04 LTS, equipped with an Intel Core i7-12700H processor (12th generation, Alder Lake hybrid architecture with 6 performance and 8 efficient cores) and 64 GB of DDR4 memory. The code was written entirely in C++17 and parallelized…