16 papers · ranked by Valyu relevance
Pieter Pas
We present Cyqlone, a solver for linear systems with a stage-wise optimal control structure that fully exploits the various levels of parallelism available in modern hardware. Cyqlone unifies algorithms based on the sequential Riccati recursion, parallel Schur complement methods, and cyclic reduction methods, thereby…
Erna Begović Kovač, Vjeran Hari
The paper analyzes special cyclic Jacobi methods for symmetric matrices of order 4. Only those cyclic pivot strategies that enable full parallelization of the method are considered. These strategies, unlike the serial pivot strategies, can force the method to be very slow or very fast within one cycle, depending on the…
Jovan Blanuša, Paolo Ienne, Kubilay Atasu
Enumerating simple cycles has important applications in computational biology, network science, and financial crime analysis. In this work, we focus on parallelising the state-of-the-art simple cycle enumeration algorithms by Johnson and Read-Tarjan along with their applications to temporal graphs. To our knowledge, we…
Qinmeng Zou, Frédéric Magoulès
This paper proposes a new gradient method to solve the large-scale problems. Theoretical analysis shows that the new method has finite termination property for two dimensions and converges R-linearly for any dimensions. Experimental results illustrate first the issue of parallel implementation. Then, the solution of a…
Vjeran Hari, Erna Begović
The paper studies the global convergence of the block Jacobi method for symmetric matrices. Given a symmetric matrix A of order n, the method generates a sequence of matrices by the rule A (k+1) = U T k A (k)Uk, k ≥ 0, where Uk are orthogonal elementary block matrices. A class of generalized serial pivot strategies is…
Gang Liao, Si-hui Qin, Longfei Ma, Qi Sun
In this paper, we focus on the need for two approaches to optimize producer and consumer synchronization for autoparallelizing compiler. Emphasis is placed on the construction of a criterion model by which the compiler reduce the number of synchronization operations needed to synchronize the dependence in a loop and…
Hang Song, Kristen Matsuno, Jacob R. West, Akshay Subramaniam + 2 more
'Aditya S. Ghate' 'Sanjiva K. Lele'] A scalable algorithm for solving compact banded linear systems on distributed memory architectures is presented. The proposed method factorizes the original system into two levels of memory hierarchies, and solves it using parallel cyclic reduction on both distributed and shared…
Meisam Razaviyayn, Mingyi Hong, Zhi‐Quan Luo, Jong‐Shi Pang
Consider the problem of minimizing the sum of a smooth (possibly non-convex) and a convex (possibly nonsmooth) function involving a large number of variables. A popular approach to solve this problem is the block coordinate descent (BCD) method whereby at each iteration only one variable block is updated while the…
Yang Yu
Sequence Authors: ['Yang Yu'] Recent advances in reasoning models have demonstrated significant improvements in accuracy, particularly for complex tasks such as mathematical reasoning, by employing detailed and comprehensive reasoning processes. However, generating these lengthy reasoning sequences is computationally…
Daniel Ruprecht
The paper introduces an OpenMP implementation of pipelined Parareal and compares it to a standard MPI-based implementation. Both versions yield essentially identical runtimes, but, depending on the compiler, the OpenMP variant consumes about 7% less energy. However, its key advantage is a significantly smaller memory…
Authors not listed
Modeling peptide cyclization is critical for the virtual screening of candidate peptides with desirable physical and pharmaceutical properties. This task is challenging because a cyclic peptide often exhibits diverse, ring-shaped conformations, which cannot be well captured by deterministic prediction models derived…
Adrian Jackson, Orestis Agathokleous
—Regions of nested loops are a common feature of High Performance Computing (HPC) codes. In shared memory programming models, such as OpenMP, these structure are the most common source of parallelism. Parallelising these structures requires the programmers to make a static decision on how parallelism should be applied.…
Donald S. Ene, V.I.E Anireh
- Evaluating how well a whole system or set of subsystems performs is one of the primary objectives of performance testing. We can tell via performance assessment if the architecture implementation meets the design objectives. Performance evaluations of several parallel algorithms are compared in this study. Both…
Temitayo Adefemi
—Parallelization has become a cornerstone of modern computing, influencing everything from high-performance supercomputers to everyday mobile devices. This paper presents a comprehensive guide on the fundamentals of parallelization that every computer scientist should know, beginning with a historical perspective that…
Vitaly Aksenov, Petr Kuznetsov, Anatoly Shalyto
A parallel batched data structure is designed to process synchronized batches of operations on the data structure using a parallel program. In this paper, we propose parallel combining, a technique that implements a concurrent data structure from a parallel batched one. The idea is that we explicitly synchronize…
V. Schwambach, S. Cleyet-Merle, Alain Issard, Stéphane Mancini
—Computer vision applications constitute one of the key drivers for embedded multicore architectures. Although the number of available cores is increasing in new architectures, designing an application to maximize the utilization of the platform is still a challenge. In this sense, parallel performance prediction tools…