12 papers · ranked by Valyu relevance
Pieter Pas
We present Cyqlone, a solver for linear systems with a stage-wise optimal control structure that fully exploits the various levels of parallelism available in modern hardware. Cyqlone unifies algorithms based on the sequential Riccati recursion, parallel Schur complement methods, and cyclic reduction methods, thereby…
Shiv Sundram, Akhilesh Balasingam, Nathan Zhang, Kunle Olukotun + 1 more
We present Cyclotron, a framework and compiler for using recurrence equations to express streaming dataflow algorithms, which then get portably compiled to distributed topologies of interlinked processors. Our framework provides an input language of recurrences over logical tensors, which then gets lowered into an…
Mohammad Javad Khani, Mahmood Ahmadi
Cyclic Redundancy Check (CRC) remains one of the most widely used error-detection mechanisms in communication, storage, and embedded systems. However, conventional software CRC implementations suffer from inherent sequential dependencies that limit efficient utilization of modern multi-core processors. This paper…
Zhixin Ou, Peng Liang, Jianchen Han, Baihui Liu + 1 more
Dynamic sequences with varying lengths have been widely used in the training of Transformer-based large language models (LLMs). However, current training frameworks adopt a pre-defined static parallel strategy for these sequences, causing neither communication-parallelization cancellation on short sequences nor…
Ramon Moya
We develop the ARE method (Action-Rectification-Expansion), a structural framework for the organization of Leibniz terms in determinants through cyclic group actions and orbital decompositions. The symmetric group S_n is partitioned into (n-1)! disjoint orbits of size n under right composition by the cyclic group C_n.…
Elwood, Alex, Tom Deakin, Justin Lovegrove + 1 more
Discrete ordinates S N transport solvers on unstructured meshes pose a challenge to scale due to complex data dependencies, memory access patterns and a highdimensional domain. In this paper, we review the performance bottlenecks within the shared memory parallelization scheme of an existing transport solver on modern…
Shiting Long, Gustavo Ramirez-Hidalgo, Andreas Frommer, Dirk Pleiter
Gauss-Seidel is a well-established iterative method for the solution of linear systems, and multicoloring has been widely used to increase parallelism in iterative solution techniques. Implementing multi-color Gauss-Seidel with conventional divide-and-conquer parallelization strategies, however, may be inefficient due…
Serge Gratton, Philippe L. Toint
A new class of asynchronous adaptive first-order optimization methods is introduced, comprising asynchronous variants of several popular algorithms. Versions of these methods using momentum and/or inexact normalization are also considered. The convergence of methods in the class on non-convex functions is analyzed in a…
Michael T. Goodrich, Vinesh Sridhar
Embedded systems and Internet of Things (IoT) applications motivate in-place parallel algorithms, which avoid allocating additional shared memory past the input. Work by Gu, Obeya, and Shun [APOCS '21] defines a family of PIP (parallel in-place) models and parallel algorithms that eschew auxiliary memory at high…
Bin Fu
We develop a parallel framework that assembles static gradient methods to achieve better adaptivity. A static gradient method, denoted by $\mathrm{GD}(x_0,T)$, takes as input an initial point $x_0\in\mathbb{R}^n$ and $T\in \mathbb{R}^+$ specifying the number $\floor{T}$ of iterations. The step size is chosen as…
Prajjwal Nijhara, Dip Sankar Banerjee
Maximal Independent Set (MIS) in a graph is a fundamental problem with applications in resource allocation, scheduling, and network optimization. Although graphs are inherently un-structured and challenging for GPU parallelism due to irregular memory access and workload imbalance, specialized GPU algorithms have…
V. Blomer, Kai-Uwe Bux
We discuss a modification to the Gries–Mills block swapping scheme for in-place rotation with average costs of 1.85 moves per element and worst case performance still at 3 moves per element. Analysis of the average case relies on the asymptotic behavior of the sum of remainders in the Euclidean algorithm.