Search · four archives
Search · four archives
12 papers · ranked by Valyu relevance
Mingjing Li, Huihui Zhou, Xiaofeng Xu, Zhiwei Zhong + 15 more
There is a growing necessity for edge training to adapt to dynamically changing environments. Neuromorphic computing represents a significant pathway for highly efficient intelligent computation in energy-constrained edges, but existing neuromorphic architectures lack the ability of directly training spiking neural…
Ran Ginosar
I have greatly enjoyed spending many years in studying parallel computing. My journey goes thorough MP-C, PLURAL, Async Plural, HAL, RC64 and more. As a PhD student at Princeton I studied a combination of shared memory and message passing, motivated by algorithms and the ease of programming. While at the Technion, a…
Shiting Long, Gustavo Ramirez-Hidalgo, Andreas Frommer, Dirk Pleiter
Gauss-Seidel is a well-established iterative method for the solution of linear systems, and multicoloring has been widely used to increase parallelism in iterative solution techniques. Implementing multi-color Gauss-Seidel with conventional divide-and-conquer parallelization strategies, however, may be inefficient due…
Aleix Roca, Vicenç Beltran
The convergence of high-performance computing (HPC) and artificial intelligence (AI) is driving the emergence of increasingly complex parallel applications and workloads. These workloads often combine multiple parallel runtimes within the same application or across co-located jobs, creating scheduling demands that…
Minyu Cheng, Jiakun Yan, Marc Snir
The bulk synchronous parallel (BSP) model struggles with irregular workloads due to rigid global communication. While fine-grained asynchronous BSP (FA-BSP) improves overlap, existing implementations typically rely on a limiting one-process-per-core model. This paper proposes a multithreaded FA-BSP approach combining…
Villalobos, Johansell, Ruzicka, Josef + 2 more
—Scientific computing in the exascale era demands increased computational power to solve complex problems across various domains. With the rise of heterogeneous computing architectures the need for vendor-agnostic, performance portability frameworks has been highlighted. Libraries like Kokkos have become essential for…
Haymo Kutschbach
This work introduces a self-optimizing virtual processor (VP) for numerical array programs that shifts parallelization from a manual developer task to a cooperative, agent-like runtime mechanism. Instead of relying on centralized task-graph scheduling, static compiler optimization, or explicitly annotated parallel…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
Background The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated…
Rajeswari Suance P S, Ruchika Gupta, Maurizio Palesi, John Jose
Multi-core quantum computing architectures have emerged as a promising solution to the qubit scalability limitations of monolithic NISQ devices. Quantum algorithms are expressed as quantum circuits composed of single- and two-qubit gates. However, circuit scheduling in multi-core quantum systems remains largely…
Kevin Garner, Chander Sadasivan, Nikos Chrisochoides
This paper presents two performance optimization techniques for a mesh adaptation method that is designed to help streamline the discretization of complex vascular geometries within the numerical modeling process. This method is integrated into a pipeline with an image-to-mesh conversion tool to generate adaptive…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…