12 papers · ranked by Valyu relevance
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Mohammad Abdur Rob, Md. Zakir Hossen, Md. Kamal Hossen, Md. Mithun Ali + 2 more
Sorting algorithms play a crucial role in computing, but most are designed with rigid structure that are only efficient under certain conditions. Although some sorting algorithms perform well in some circumstances, they do not perform well on some resistant platforms. This study introduces Wall-L Merge Sort, which…
Sean O’Connell, Jonathan A. Michaels, Runming Wang, Sahit Mamidipaka + 6 more
Understanding how neural signals control muscle activity during behavior is a key challenge in motor neuroscience. To this end, recent advances in intramuscular multielectrode arrays have enabled high-quality multichannel recordings of many motor unit action potentials (MUAPs) in freely moving subjects. However…
Alessio Paolo Buccino, Arjun Sridhar, David Feng, Karel Svoboda + 3 more
The scale of in vivo electrophysiology has expanded in recent years, with simultaneous recordings across thousands of electrodes now becoming routine. These advances have enabled a wide range of discoveries, but they also impose substantial computational demands. Spike sorting, the procedure that extracts spikes from…
Shujun Peng, Xinhan Lin, Yu Zhang, Yuheng Xiao + 1 more
Parallel scan is a fundamental primitive widely used in a broad range of workloads, including parallel sorting, graph algorithms, and sampling in large language model inference. Although GPU-optimized parallel scan algorithms have been extensively studied, their reliance on vector units makes them inefficient on modern…
Zhonghai Zhang, Yewen Li, Ke Meng, Chunming Zhang + 2 more
FastDup is implemented in C++ and leverages the htslib () for high-performance parsing of SAM/BAM files. To maximize I/O efficiency, it uses an asynchronous I/O pipeline that decouples file decompression/compression from computational tasks, enabling simultaneous multi-threaded read and write operations through…
Kecong Tang, Ardalan Naseri, Degui Zhi, Shaojie Zhang + 1 more
At each site indexed by k, the algorithm maintains a prefix array $P_{k}$, which encodes the co-lexicographic ordering of haplotypes based on their reversed prefixes up to the current site. Formally, let ${x_{i}}$, where $i=1$ to M, denote M haplotype sequences over N sites. At site k, the prefix array $P_{k}$ is…
Julia Golonka, Filip Krużel, Rosario Schiano Lo Moriello
Resource-constrained sensor nodes in Internet-of-Things (IoT) and embedded sensing applications frequently rely on low-cost microcontrollers, where even basic algorithmic choices directly impact latency, energy consumption, and memory footprint. This study evaluates six sorting algorithms-Bubble Sort, Insertion Sort…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
Christoph Stelz, Lukas Hübner, Alexandros Stamatakis, Russell Schwartz
Given an input array of elements $E=[e_{0},e_{1},…,e_{n-1}]$, distributed across p processing elements (PEs; e.g. processes or threads), we desire to compute $r=e_{0}⊕e_{1}⊕…⊕e_{n-1}$, where $⊕$ denotes a binary, associative operation (e.g. summation or multiplication). In a distributed reduction, we return the result…
Sangjin Lee, Sunggon Kim, Yongseok Son, Agbotiname Lucky Imoize
We propose ScaleDefrag, a parallel and asynchronous defragmentation tool that reduces defragmentation time by up to 3.8× compared to e4defrag, while improving scalability on multi-core systems. Flash-based solid-state drives (SSDs) have been widely adopted in various large-scale storage systems including cloud and HPC…
Mateusz Gruzewski, Marek Palkowski, Ramon Antonio Rodriges Zalipynis
In this article, we present an efficient and concise OpenMP implementation of the Nussinov RNA folding algorithm, a well-known representative of non-serial polyadic dynamic programming (NPDP). Our goal is to develop an optimized implementation that can serve as a template for related dynamic programming applications.…