11 papers · ranked by Valyu relevance
Michael T. Goodrich, Vinesh Sridhar
Embedded systems and Internet of Things (IoT) applications motivate in-place parallel algorithms, which avoid allocating additional shared memory past the input. Work by Gu, Obeya, and Shun [APOCS '21] defines a family of PIP (parallel in-place) models and parallel algorithms that eschew auxiliary memory at high…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Weitian Chen, Shixuan Sun, Cheng Chen, Yongmin Hu + 2 more
Subgraph matching is a core operation in graph analytics, supporting a broad spectrum of applications from social network analysis to bioinformatics. Recent GPU-based approaches accelerate subgraph matching by leveraging parallelism but rely on a coarse-grained execution model that suffers from scalability and…
Zhixin Ou, Peng Liang, Jianchen Han, Baihui Liu + 1 more
Dynamic sequences with varying lengths have been widely used in the training of Transformer-based large language models (LLMs). However, current training frameworks adopt a pre-defined static parallel strategy for these sequences, causing neither communication-parallelization cancellation on short sequences nor…
Amir Hossein Salehi Shayegan
In this work, we present a solution to the critical limitation of qubit capacity in near-term quantum hardware by giving a hybrid framework that integrates the spectral element method (SEM) with distributed quantum computing. Using domain decomposition techniques, the additive and multiplicative Schwarz methods, the…
Shujun Peng, Xinhan Lin, Yu Zhang, Yuheng Xiao + 1 more
Parallel scan is a fundamental primitive widely used in a broad range of workloads, including parallel sorting, graph algorithms, and sampling in large language model inference. Although GPU-optimized parallel scan algorithms have been extensively studied, their reliance on vector units makes them inefficient on modern…
Accorsi, Luca, Laganà, Demetrio + 6 more
We propose a parallel shared-memory schema to cooperatively optimize the solution of a Capacitated Vehicle Routing Problem instance with minimal synchronization effort and without the need for an explicit decomposition. To this end, we design FILO2 x as a single-trajectory parallel adaptation of the FILO2 algorithm…
Mohammad Abdur Rob, Md. Zakir Hossen, Md. Kamal Hossen, Md. Mithun Ali + 2 more
Sorting algorithms play a crucial role in computing, but most are designed with rigid structure that are only efficient under certain conditions. Although some sorting algorithms perform well in some circumstances, they do not perform well on some resistant platforms. This study introduces Wall-L Merge Sort, which…
Teh Noranis Mohd Aris, Ningning Chen, Norwati Mustapha, Maslina Zolkepli + 1 more
To address the inefficiencies in sample utilization and policy instability in asynchronous distributed reinforcement learning, we propose TPDEB-a dual experience replay framework that integrates prioritized sampling and temporal diversity. While recent distributed RL systems have scaled well, they often suffer from…
Marco Ronzani, Cristina Silvano
Hypergraph partitioning is a pervasive NP-hard problem, and accelerating its computation on GPU can both slice time-to-solution and raise quality of results. In this work, we implement a multi-level hypergraph partitioning algorithm on GPU targeting a specific set of problem constraints: bounded per-partition size and…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…