13 papers · ranked by Valyu relevance
Ji Liu, Stephen J. Wright
The randomized Kaczmarz (RK) algorithm is a simple but powerful approach for solving consistent linear systems Ax = b. This paper proposes an accelerated randomized Kaczmarz (ARK) algorithm with better convergence than the standard RK algorithm on ill conditioned problems. The per-iteration cost of RK and ARK are…
Gang Mei, Nengxiong Xu, Liangliang Xu
This paper presents an efficient parallel Adaptive Inverse Distance Weighting (AIDW) interpolation algorithm on modern Graphics Processing Unit (GPU). The presented algorithm is an improvement of our previous GPU-accelerated AIDW algorithm by adopting fast k-nearest neighbors (kNN) search. In AIDW, it needs to find…
Xuan Zuo, Hui-Yan Li, Shan Gao, Pu Zhang + 2 more
Adaptive gradient algorithms have been successfully used in deep learning. Previous work reveals that adaptive gradient algorithms mainly borrow the moving average idea of heavy ball acceleration to estimate the first- and second-order moments of the gradient for accelerating convergence. However, Nesterov acceleration…
Michael Muehlebach, Michael I. Jordan
We exploit analogies between first-order algorithms for constrained optimization and non-smooth dynamical systems to design a new class of accelerated first-order algorithms for constrained optimization. Unlike Frank-Wolfe or projected gradients, these algorithms avoid optimization over the entire feasible set at each…
Palma London, Shai Vardi, Adam Wierman, Hanling Yi
This paper presents an acceleration framework for packing linear programming problems where the amount of data available is limited, i.e., where the number of constraints m is small compared to the variable dimension n. The framework can be used as a black box to speed up linear programming solvers dramatically, by two…
Yu-Wen Chen, Antonio Orvieto, Aurélien Lucchi
Derivative-free optimization (DFO) has recently gained a lot of momentum in machine learning, spawning interest in the community to design faster methods for problems where gradients are not accessible. While some attention has been given to the concept of acceleration in the DFO literature, existing stochastic…
Soheil Shahrouz, Saber Salehkaleybar, Matin Hashemi
—Given a social network modeled as a weighted graph G, the influence maximization problem seeks k vertices to become initially influenced, to maximize the expected number of influenced nodes under a particular diffusion model. The influence maximization problem has been proven to be NP-hard, and most proposed solutions…
Xin Wang, Bin Zhang, Xu Cao, Fei Liu + 2 more
Fluorescence molecular tomography (FMT) with early-photons can improve the spatial resolution and fidelity of the reconstructed results. However, its computing scale is always large which limits its applications. In this paper, we introduced an acceleration strategy for the early-photon FMT with graphics processing…
Yangyang Xu
Motivated by big data applications, first-order methods have been extremely popular in recent years. However, naive gradient methods generally converge slowly. Hence, much efforts have been made to accelerate various first-order methods. This paper proposes two accelerated methods towards solving structured linearly…
Pranay Reddy Kommera, Vinay Ramakrishnaiah, Christine Sweeney, Jeffrey Donatelli + 1 more
'Jeffrey Donatelli' 'Petrus H. Zwart'] The paper presents efforts to accelerate the multitiered iterative phasing (MTIP) algorithm on contemporary graphics processing units (GPUs). Application portability is demonstrated by accelerating the MTIP algorithm on NVIDIA and AMD GPUs using a single codebase.
Zixuan Li, Mingxing Duan, Huizhang Luo, Wangdong Yang + 2 more
Using GPU Tensor Cores Authors: ['Zixuan Li' 'Mingxing Duan' 'Huizhang Luo' 'Wangdong Yang' 'Kenli Li' 'Keqin Li'] Abstract—Sparse tensors are prevalent in real-world applications, often characterized by their large-scale, high-order, and highdimensional nature. Directly handling raw tensors is impractical due to the…
Shubhendra Pal Singhal, M. Srıdevı
—Optimization of searching the best possible action depending on various states like state of environment, system goal etc. has been a major area of study in computer systems. In any search algorithm, searching best possible solution from the pool of every possibility known can lead to the construction of the whole…
Jonas Latt, Christophe Coreixas, Joël Beny, Fang-Bao Tian
We present a novel, hardware-agnostic implementation strategy for lattice Boltzmann (LB) simulations, which yields massive performance on homogeneous and heterogeneous many-core platforms. Based solely on C++17 Parallel Algorithms, our approach does not rely on any language extensions, external libraries…