14 papers · ranked by Valyu relevance
Yiling Xie, Yiling Luo, Xiaoming Huo
A primal-dual accelerated stochastic gradient descent with variance reduction algorithm (PDASGD) is proposed to solve linear-constrained optimization problems. PDASGD could be applied to solve the discrete optimal transport (OT) problem and enjoys the best-known computational complexity—Oe(n 2/ϵ), where n is the number…
Chang He, Zhaoye Pan, Xiao Wang, Bo Jiang
Lower Query Complexity Authors: ['Chang He' 'Zhaoye Pan' 'Xiao Wang' 'Bo Jiang'] Optimization problems with access to only zerothorder information of the objective function on Riemannian manifolds arise in various applications, spanning from statistical learning to robot learning. While various zeroth-order algorithms…
Roberto Carrasco, Enzo Meneses, Héctor Ferrada, Cristóbal A. Navarro + 1 more
In recent years, applications such as real-time simulations, autonomous systems, and video games increasingly demand the processing of complex geometric models under stringent time constraints. Traditional geometric algorithms, including the convex hull, are subject to these challenges. A common approach to improve…
Marco Angioli, Marcello Barbirotta, Abdallah Cheikh, Antonio Mastrandrea + 2 more
and Vector Computing Acceleration for Embedded Learning Systems Authors: ['Marco Angioli' 'Marcello Barbirotta' 'Abdallah Cheikh' 'Antonio Mastrandrea' 'Francesco Menichelli' 'Mauro Olivieri'] As the Internet of Things expands, embedding Artificial Intelligence algorithms in resource-constrained devices has become…
Zebang Shen, Hui Qian, Tongzhou Mu, Chao Zhang
Nowadays, algorithms with fast convergence, small memory footprints, and low per-iteration complexity are particularly favorable for artificial intelligence applications. In this paper, we propose a doubly stochastic algorithm with a novel accelerating multi-momentum technique to solve large scale empirical risk…
Xiaochuan Gong, Jie Hao, Mingrui Liu
Unbounded Smoothness Authors: ['Xiaochuan Gong' 'Jie Hao' 'Mingrui Liu'] This paper investigates a class of stochastic bilevel optimization problems where the upper-level function is nonconvex with potentially unbounded smoothness and the lower-level problem is strongly convex. These problems have significant…
Ya‐Nan Zhu
This work proposes an Accelerated Primal–Dual Fixed-Point (APDFP) method that employs Nesterov type acceleration to solve composite problems of the form min x f(x) + g ◦ B(x), where g is nonsmooth and B is a linear operator. The APDFP features fully decoupled iterations and can be regarded as a generalization of…
Guojin Chen, Haoyu Yang, Bei Yu
Multiple patterning lithography (MPL) is regarded as one of the most promising ways of overcoming the resolution limitations of conventional optical lithography due to the delay of next-generation lithography technology. As the feature size continues to decrease, layout decomposition for multiple patterning lithography…
Kewang Chen, Ye Ji, Matthias Möller, C. Vuik
In this report, we present a versatile and efficient preconditioned Anderson acceleration (PAA) method for fixed-point iterations. The proposed framework offers flexibility in balancing convergence rates (linear, super-linear, or quadratic) and computational costs related to the Jacobian matrix. Our approach recovers…
Zixuan Li, Mingxing Duan, Huizhang Luo, Wangdong Yang + 2 more
Using GPU Tensor Cores Authors: ['Zixuan Li' 'Mingxing Duan' 'Huizhang Luo' 'Wangdong Yang' 'Kenli Li' 'Keqin Li'] Abstract—Sparse tensors are prevalent in real-world applications, often characterized by their large-scale, high-order, and highdimensional nature. Directly handling raw tensors is impractical due to the…
Shan Haoxuan, Guo, Cong, Wei + 6 more
—The rapid scaling of large language models demands more efficient hardware. Quantization offers a promising trade-off between efficiency and performance. With ultra-low-bit quantization, there are abundant opportunities for results reuse, and thus it can be boosted with lookup tables (LUTs) based acceleration.…
Hassan Nassar, Rafik Youssef, Lars Bauer, Jörg Henkel
As the need for more computing power grows, traditional methods are hitting limits. To boost performance, we're expanding Central Processing Unit (CPU) capabilities and using specialized hardware accelerators. For example, mobile devices usually have cameras, video encoding, and audio accelerators. To perform the…
Qian Chen, Yang Xiao-feng, Shengli Lu
for SpTRSV Authors: ['Qian Chen' 'Yang Xiao-feng' 'Shengli Lu'] Abstract—Sparse triangular solve (SpTRSV) is widely used in various domains. Numerous studies have been conducted using CPUs, GPUs, and specific hardware accelerators, where dataflow can be categorized into coarse and fine granularity. Coarse dataflow…
Gangli Liu
problem in an undirected dense graph Authors: ['Gangli Liu'] We provide an efficient ( 2 ) implementation for solving the all pairs minimax path problem or widest path problem in an undirected dense graph. It is a code implementation of the Algorithm 4 (MMJ distance by Calculation and Copy) in a previous paper. The…