15 papers · ranked by Valyu relevance
Shridharan Chandramouli
| 1 | | Introduction | | | |---|-----------------------------------------------------|---------------------------------------------------------------------------|----|--| | | 1.1 | Motivation | 1 | | | | 1.2 | Proper Order Multi-Column Graph Structure | 2 | | | | 1.3 | General Maxflow/Mincut Problem Definition | 2 | |…
Simo Särkkä, Ángel F. García‐Fernández
—This paper presents an experimental evaluation of parallel-in-time Kalman filters and smoothers using graphics processing units (GPUs). In particular, the paper evaluates different all-prefix-sum algorithms, that is, parallel scan algorithms for temporal parallelization of Kalman filters and smoothers in two ways: by…
Michael T. Goodrich, Vinesh Sridhar
Embedded systems and Internet of Things (IoT) applications motivate in-place parallel algorithms, which avoid allocating additional shared memory past the input. Work by Gu, Obeya, and Shun [APOCS '21] defines a family of PIP (parallel in-place) models and parallel algorithms that eschew auxiliary memory at high…
Weitian Chen, Shixuan Sun, Cheng Chen, Yongmin Hu + 2 more
Subgraph matching is a core operation in graph analytics, supporting a broad spectrum of applications from social network analysis to bioinformatics. Recent GPU-based approaches accelerate subgraph matching by leveraging parallelism but rely on a coarse-grained execution model that suffers from scalability and…
Zhixin Ou, Peng Liang, Jianchen Han, Baihui Liu + 1 more
Dynamic sequences with varying lengths have been widely used in the training of Transformer-based large language models (LLMs). However, current training frameworks adopt a pre-defined static parallel strategy for these sequences, causing neither communication-parallelization cancellation on short sequences nor…
Hutton, Chase, Melrod, Adam
We present the first parallel batch-dynamic algorithm for maintaining a proper (Δ + 1)-vertex coloring. Our approach builds on a new sequential dynamic algorithm inspired by the work of Bhattacharya et al. (SODA'18). The resulting randomized algorithm achieves (log Δ) expected amortized update time and, for any batch…
Accorsi, Luca, Laganà, Demetrio + 6 more
We propose a parallel shared-memory schema to cooperatively optimize the solution of a Capacitated Vehicle Routing Problem instance with minimal synchronization effort and without the need for an explicit decomposition. To this end, we design FILO2 x as a single-trajectory parallel adaptation of the FILO2 algorithm…
Ziyang Men, Bo Huang, Yan Gu, Yihan Sun
Maintaining spatial data (points in two or three dimensions) is crucial and has a wide range of applications, such as graphics, GIS, and robotics. To support efficient updates and queries on the spatial data, many data structures, called spatial indexes, have been proposed, e.g., d-trees, oct/quadtrees (also called…
Wolfgang Bangerth
Digital elevation models (DEMs) have reached resolutions and sizes that only parallel computaters can efficiently process. One important application of DEMs is predicting how much water flows where, the so-called ``flow routing problem'' (a variation of which is the problem of determining the drainage area upstream of…
Ajay Singh, Nikos Metaxakis, Panagiota Fatourou
We present a new blocking linearizable stack implementation which utilizes sharding and fetch&increment to achieve significantly better performance than all existing concurrent stacks. The proposed implementation is based on a novel elimination mechanism and a new combining approach that are efficiently blended to gain…
Shruthi Kannappan, Ashwina Kumar, Rupesh Nasre
MaxFlow is a fundamental problem in graph theory and combinatorial optimisation, used to determine the maximum flow from a source node to a sink node in a flow network. It finds applications in diverse domains, including computer networks, transportation, and image segmentation. The core idea is to maximise the total…
Carl Kugblenu, Petri Vuorimaa
Production vector search systems often fan out each query across parallel lanes (threads, replicas, or shards) to meet latency service-level objectives (SLOs). In practice, these lanes rediscover the same candidates, so extra compute does not increase coverage. We present a coordination-free lane partitioner that turns…
Minyu Cheng, Jiakun Yan, Marc Snir
The bulk synchronous parallel (BSP) model struggles with irregular workloads due to rigid global communication. While fine-grained asynchronous BSP (FA-BSP) improves overlap, existing implementations typically rely on a limiting one-process-per-core model. This paper proposes a multithreaded FA-BSP approach combining…
Michał Szyfelbein
Consider the classical \textsc{Min-Sum Set Cover} problem: We are given a universe $\mathcal{U}$ of $n$ elements and a collection $\mathcal{S}$ of $k$ subsets of $\mathcal{U}$. Moreover, a cost function is associated with each set. The goal is to find a subsequence of sets in $\mathcal{S}$ that covers all elements in…
Marco Ronzani, Cristina Silvano
Hypergraph partitioning is a pervasive NP-hard problem, and accelerating its computation on GPU can both slice time-to-solution and raise quality of results. In this work, we implement a multi-level hypergraph partitioning algorithm on GPU targeting a specific set of problem constraints: bounded per-partition size and…