13 papers · ranked by Valyu relevance
Michael T. Goodrich, Vinesh Sridhar
Embedded systems and Internet of Things (IoT) applications motivate in-place parallel algorithms, which avoid allocating additional shared memory past the input. Work by Gu, Obeya, and Shun [APOCS '21] defines a family of PIP (parallel in-place) models and parallel algorithms that eschew auxiliary memory at high…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Amir Hossein Salehi Shayegan
In this work, we present a solution to the critical limitation of qubit capacity in near-term quantum hardware by giving a hybrid framework that integrates the spectral element method (SEM) with distributed quantum computing. Using domain decomposition techniques, the additive and multiplicative Schwarz methods, the…
Zhixin Ou, Peng Liang, Jianchen Han, Baihui Liu + 1 more
Dynamic sequences with varying lengths have been widely used in the training of Transformer-based large language models (LLMs). However, current training frameworks adopt a pre-defined static parallel strategy for these sequences, causing neither communication-parallelization cancellation on short sequences nor…
Mohammad Abdur Rob, Md. Zakir Hossen, Md. Kamal Hossen, Md. Mithun Ali + 2 more
Sorting algorithms play a crucial role in computing, but most are designed with rigid structure that are only efficient under certain conditions. Although some sorting algorithms perform well in some circumstances, they do not perform well on some resistant platforms. This study introduces Wall-L Merge Sort, which…
Rubén Langarita, Jesús Alastruey-Benedé, Pablo Ibáñez, Santiago Marco‐Sola + 2 more
—Multiple HPC applications are often bottlenecked by compute-intensive kernels implementing complex dependency patterns (data-dependency bound). Traditional general-purpose accelerators struggle to effectively exploit fine-grain parallelism due to limitations in implementing convoluted data-dependency patterns (like…
Accorsi, Luca, Laganà, Demetrio + 6 more
We propose a parallel shared-memory schema to cooperatively optimize the solution of a Capacitated Vehicle Routing Problem instance with minimal synchronization effort and without the need for an explicit decomposition. To this end, we design FILO2 x as a single-trajectory parallel adaptation of the FILO2 algorithm…
Weitian Chen, Shixuan Sun, Cheng Chen, Yongmin Hu + 2 more
Subgraph matching is a core operation in graph analytics, supporting a broad spectrum of applications from social network analysis to bioinformatics. Recent GPU-based approaches accelerate subgraph matching by leveraging parallelism but rely on a coarse-grained execution model that suffers from scalability and…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
Lorién López-Villellas, Cristian Iñiguez, Albert Jiménez-Blanco, Quim Aguado-Puig + 5 more
Since each core performs the same workload as in the single-threaded benchmarks, total memory usage at a given thread count can be precisely computed by multiplying the per-thread memory reported in [btag183-T2] by the number of threads. The resulting multi-threaded memory usage is provided in the [sup1], available as…
Marc Becker, Bernd Bischl
Many algorithms in statistics and machine learning can be parallelized in an asynchronous manner where workers need to communicate through shared state rather than execute independent tasks dispatched by a central controller. Especially in modern hyperparameter optimization and parallel black-box optimization with…
Christoph Stelz, Lukas Hübner, Alexandros Stamatakis, Russell Schwartz
Given an input array of elements $E=[e_{0},e_{1},…,e_{n-1}]$, distributed across p processing elements (PEs; e.g. processes or threads), we desire to compute $r=e_{0}⊕e_{1}⊕…⊕e_{n-1}$, where $⊕$ denotes a binary, associative operation (e.g. summation or multiplication). In a distributed reduction, we return the result…
Shiting Long, Gustavo Ramirez-Hidalgo, Andreas Frommer, Dirk Pleiter
Gauss-Seidel is a well-established iterative method for the solution of linear systems, and multicoloring has been widely used to increase parallelism in iterative solution techniques. Implementing multi-color Gauss-Seidel with conventional divide-and-conquer parallelization strategies, however, may be inefficient due…