12 papers · ranked by Valyu relevance
de Ronde, Folkert, Knapen, Alexander + 4 more
This paper introduces a novel approach to enhance the execution of quantum algorithms on distributed quantum systems. The proposed method involves the development of a hardware design that supports parallel instruction execution and a compiler that modifies the order of instructions to increase parallelism…
Haymo Kutschbach
This work introduces a self-optimizing virtual processor (VP) for numerical array programs that shifts parallelization from a manual developer task to a cooperative, agent-like runtime mechanism. Instead of relying on centralized task-graph scheduling, static compiler optimization, or explicitly annotated parallel…
Shiting Long, Gustavo Ramirez-Hidalgo, Andreas Frommer, Dirk Pleiter
Gauss-Seidel is a well-established iterative method for the solution of linear systems, and multicoloring has been widely used to increase parallelism in iterative solution techniques. Implementing multi-color Gauss-Seidel with conventional divide-and-conquer parallelization strategies, however, may be inefficient due…
Aleix Roca, Vicenç Beltran
The convergence of high-performance computing (HPC) and artificial intelligence (AI) is driving the emergence of increasingly complex parallel applications and workloads. These workloads often combine multiple parallel runtimes within the same application or across co-located jobs, creating scheduling demands that…
Elwood, Alex, Tom Deakin, Justin Lovegrove + 1 more
Discrete ordinates S N transport solvers on unstructured meshes pose a challenge to scale due to complex data dependencies, memory access patterns and a highdimensional domain. In this paper, we review the performance bottlenecks within the shared memory parallelization scheme of an existing transport solver on modern…
Harry Fitchett, Charles M. Fox
While transistor density is still increasing, clock speeds are not, motivating the search for new parallel architectures. One approach is to completely abandon the concept of CPU – and thus serial imperative programming – and instead to specify and execute tasks in parallel, compiling from programming languages to data…
Harry Fitchett, Jasmine Ritchie, Charles M. Fox
Harry Fitchett, Jasmine Ritchie, Charles Fox, School of Engineering and Physical Science,, University of Lincoln, UK. 2026. Extending CPU-less parallel execution of lambda calculus in digital logic with lists and arithmetic. In Proceedings of Conference Name (Submission to ACM Symposium on Parallelism in Algorithms and…
Paulo Henrique Leme Ramalho, Dennis Alves Pedersen, Fábio Andrijauskas
The complexity of biomolecular simulations has substantially increased the demand for High-Performance Computing (HPC) infrastructures, particularly in molecular dynamics and coarse-grained modeling. This work presents a systematic performance and scalability analysis of the LAMMPS simulator for coarse-grained…
Alexander Strack, Alexander Van Craen, Dirk Pflüger
Fork-join parallelism, popularized by OpenMP, remains the dominant model for shared-memory parallel programming, but its implicit synchronization barriers can penalize algorithms with inhomogeneous workloads. Asynchronous many-task (AMT) runtimes sidestep these barriers by expressing work as a dependency graph of…
Matthew Diaz, Masoud Mohammadi-Arzanagh, Yingyue Zhu, Mohammad Hafezi + 3 more
Parallel processing of information plays a critical role in accelerating computation. This includes quantum computers, where parallel processing of quantum information will play a critical role in practical quantum advantage. Here, we demonstrate a new type of parallel entangling gates in a trapped-ion quantum…
Mejgan Dedaj, Argyro Gailla, Theofanis Ioannou, Stamatia Kastrinaki + 8 more
This study assesses the scalability of process-based and thread-based schedulers for many-core shared-memory systems using a memory-intensive row-wise quick-sort workload on large three-dimensional tensors. The process-based evaluation considers bounded prolific, bounded collective, and three pipe-based…
Villalobos, Johansell, Ruzicka, Josef + 2 more
—Scientific computing in the exascale era demands increased computational power to solve complex problems across various domains. With the rise of heterogeneous computing architectures the need for vendor-agnostic, performance portability frameworks has been highlighted. Libraries like Kokkos have become essential for…