15 papers · ranked by Valyu relevance
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
Background The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated…
Dipti Singh, Neha Chand
This study introduces a novel Parallel Grasshopper Optimization Algorithm (p-GOA), specifically designed to address reliability optimization problems. Although several hybrid algorithms exist in this field, the proposed p-GOA distinctly differs through its parallel cooperative strategy. Unlike sequential methods that…
Tram Nguyen, Snasel Vaclav, Bay Vo, Van Du Nguyen + 1 more
This paper proposes a parallel hybrid metaheuristic, named PH-SHOWOA, that integrates the Spotted Hyena Optimizer (SHO) and the Whale Optimization Algorithm (WOA) to solve the Vehicle Routing Problem with Simultaneous Pickup and Delivery and Time Windows (VRPSPDTW). The proposed method leverages the strength of both…
Amir Hossein Salehi Shayegan
In this work, we present a solution to the critical limitation of qubit capacity in near-term quantum hardware by giving a hybrid framework that integrates the spectral element method (SEM) with distributed quantum computing. Using domain decomposition techniques, the additive and multiplicative Schwarz methods, the…
Shujun Peng, Xinhan Lin, Yu Zhang, Yuheng Xiao + 1 more
Parallel scan is a fundamental primitive widely used in a broad range of workloads, including parallel sorting, graph algorithms, and sampling in large language model inference. Although GPU-optimized parallel scan algorithms have been extensively studied, their reliance on vector units makes them inefficient on modern…
Beenish Gul, Maria Murach, Stefan Bekiranov, Kevin Skadron + 1 more
The GPU architecture is composed of a scalable array of Streaming Multiprocessors (SMs). Modern GPUs, for example, NVIDIA’s Ampere architecture, consists of 6912 CUDA cores (NVIDIA) sliced into 108 Streaming Processors. To parallelize Leiden on GPUs, all graph data arrays, including neighbors, community assignments…
Joakim da Silva, Daniel Hernández Escobar, Tor Kjellsson Lindblom, Håkan Nordström + 1 more
$(4c)y_{1}^{(k+1)}y_{2}^{(k+1)}=(y_{1}^{k}+x_{1}^{(k+1)}-z_{1}^{(k+1)})(y_{2}^{k}+x_{2}^{(k+1)}-z_{2}^{(k+1)}),$ where the superscripts denote the iteration number and are the scaled dual variables, which are associated with an augmented Lagrangian parameter controlling the trade-off between primal and dual feasibility…
Mohammad Abdur Rob, Md. Zakir Hossen, Md. Kamal Hossen, Md. Mithun Ali + 2 more
Sorting algorithms play a crucial role in computing, but most are designed with rigid structure that are only efficient under certain conditions. Although some sorting algorithms perform well in some circumstances, they do not perform well on some resistant platforms. This study introduces Wall-L Merge Sort, which…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
Shengtian Zhang, Haolin Yang, Hyeonseok Kim, Incheol Shin + 2 more
Edge computing (EC) in the Internet of Ships (IoS) reduces the latency and energy burdens of cloud-centric architectures, but fully realizing its benefits requires effective computation offloading strategies. Designing such strategies in dynamic maritime environments remains challenging due to the high-dimensional…
Christoph Stelz, Lukas Hübner, Alexandros Stamatakis, Russell Schwartz
Given an input array of elements $E=[e_{0},e_{1},…,e_{n-1}]$, distributed across p processing elements (PEs; e.g. processes or threads), we desire to compute $r=e_{0}⊕e_{1}⊕…⊕e_{n-1}$, where $⊕$ denotes a binary, associative operation (e.g. summation or multiplication). In a distributed reduction, we return the result…
Jinyang Wang, Zhugang Wang, Di Liu, Isabelle Ledoux-Rak
Ultra-high sampling rates in coherent optical front-ends increasingly exceed the processing capabilities of real-time baseband processors, creating a bottleneck in coherent free-space optical communication systems. We propose a unified state-space framework to systematically parallelize digital signal processing (DSP)…
Mateusz Gruzewski, Marek Palkowski, Ramon Antonio Rodriges Zalipynis
In this article, we present an efficient and concise OpenMP implementation of the Nussinov RNA folding algorithm, a well-known representative of non-serial polyadic dynamic programming (NPDP). Our goal is to develop an optimized implementation that can serve as a template for related dynamic programming applications.…
Teh Noranis Mohd Aris, Ningning Chen, Norwati Mustapha, Maslina Zolkepli + 1 more
To address the inefficiencies in sample utilization and policy instability in asynchronous distributed reinforcement learning, we propose TPDEB-a dual experience replay framework that integrates prioritized sampling and temporal diversity. While recent distributed RL systems have scaled well, they often suffer from…