24 papers · ranked by Valyu relevance
Peter L. Bartlett, Chris Junchi Li, Jingfeng Wu, Bin Yu
In the field of optimization, developing accelerated methods for solving minimax and fixed-point problems remains a fundamental challenge. This paper presents a novel family of dual accelerated algorithms that achieve optimal convergence rates for both minimax and fixed-point problems. By exploring new anchoring…
Yaniv Swiel, Jean-Tristan Brandenburg, Mahtaab Hayat, Wenlong Carl Chen + 2 more
Genome-wide association studies (GWASs) analyse genetic variation over the genomes of many individuals in an attempt to identify single nucleotide polymorphisms (SNPs) associated with complex phenotypes. To capture a large amount of genetic variation and increase the chance of detecting associated SNPs, modern GWASs…
Anders Pitman, Cathy Yang, Yi Qiao
Next-generation sequencing now produces whole-genome data in hours, but downstream variant calling remains a multi-hour to multi-day bottleneck that excludes genomic analysis from time-critical clinical settings. GPU acceleration offers a natural path forward — variant calling is inherently parallelizable across…
Lingyi Chen, Haoran Tang, Hao Wu, Huihui Wu + 3 more
Numerical computation of the rate-distortion (RD) function is a key problem in RD theory. Thus far, efficient algorithms have been well studied for discrete sources, but for continuous sources, there is still lack of a rigorously developed solution. In this article, an integrated approach is conducted that bridges RD…
Chang He, Zhaoye Pan, Xiao Wang, Bo Jiang
Lower Query Complexity Authors: ['Chang He' 'Zhaoye Pan' 'Xiao Wang' 'Bo Jiang'] Optimization problems with access to only zerothorder information of the objective function on Riemannian manifolds arise in various applications, spanning from statistical learning to robot learning. While various zeroth-order algorithms…
Xuan Zuo, Hui-Yan Li, Shan Gao, Pu Zhang + 2 more
Adaptive gradient algorithms have been successfully used in deep learning. Previous work reveals that adaptive gradient algorithms mainly borrow the moving average idea of heavy ball acceleration to estimate the first- and second-order moments of the gradient for accelerating convergence. However, Nesterov acceleration…
Yiling Xie, Yiling Luo, Xiaoming Huo
A primal-dual accelerated stochastic gradient descent with variance reduction algorithm (PDASGD) is proposed to solve linear-constrained optimization problems. PDASGD could be applied to solve the discrete optimal transport (OT) problem and enjoys the best-known computational complexity—Oe(n 2/ϵ), where n is the number…
Ava Yektaeian Vaziri, Bahador Makkiabadi
This paper illustrates the development of two efficient source localization algorithms for electroencephalography (EEG) data, aimed at enhancing real-time brain signal reconstruction while addressing the computational challenges of traditional methods. Accurate EEG source localization is crucial for applications in…
Michael Muehlebach, Michael I. Jordan
We exploit analogies between first-order algorithms for constrained optimization and non-smooth dynamical systems to design a new class of accelerated first-order algorithms for constrained optimization. Unlike Frank-Wolfe or projected gradients, these algorithms avoid optimization over the entire feasible set at each…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated quadratic…
Tim Anderson, Travis J. Wheeler
Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). The Forward algorithm that provides much of…
Gaogao Liu, Wenbo Yang, Peng Li, Guodong Qin + 6 more
'Youming Wang' 'Shuai Wang' 'Ning Yue' 'Dongjie Huang' 'Luís Castedo Ribas'] The data volume and computation task of MIMO radar is huge; a very high-speed computation is necessary for its real-time processing. In this paper, we mainly study the time division MIMO radar signal processing flow, propose an improved MIMO…
Zixuan Li, Mingxing Duan, Huizhang Luo, Wangdong Yang + 2 more
Using GPU Tensor Cores Authors: ['Zixuan Li' 'Mingxing Duan' 'Huizhang Luo' 'Wangdong Yang' 'Kenli Li' 'Keqin Li'] Abstract—Sparse tensors are prevalent in real-world applications, often characterized by their large-scale, high-order, and highdimensional nature. Directly handling raw tensors is impractical due to the…
Authors not listed
Stochastic Simulation Algorithms (SSA) are a cornerstone in simulating Free Radical Polymerization (FRP) due to their accuracy and reliability. However, computational inefficiency remains a challenge for large-scale and complex polymerization systems. This work introduces a novel stochastic simulation algorithm…
Roberto Carrasco, Enzo Meneses, Héctor Ferrada, Cristóbal A. Navarro + 1 more
In recent years, applications such as real-time simulations, autonomous systems, and video games increasingly demand the processing of complex geometric models under stringent time constraints. Traditional geometric algorithms, including the convex hull, are subject to these challenges. A common approach to improve…
Kisaru Liyanage, Hiruna Samarakoon, Sri Parameswaran, Hasindu Gamaarachchi
minimap2 is the gold-standard software for reference-based sequence mapping in third-generation long-read sequencing. While minimap2 is relatively fast, further speedup is desirable, especially when processing a multitude of large datasets. In this work, we present minimap2-fpga, a hardware-accelerated version of…
Authors not listed
Modeling multimetallic systems efficiently enables faster prediction of desirable chemical properties and design of new materials. This work describes an initial implementation for performing multireference wave function method localized active space self-consistent field (LASSCF) calculations through the use of…
Hassan Nassar, Rafik Youssef, Lars Bauer, Jörg Henkel
As the need for more computing power grows, traditional methods are hitting limits. To boost performance, we're expanding Central Processing Unit (CPU) capabilities and using specialized hardware accelerators. For example, mobile devices usually have cameras, video encoding, and audio accelerators. To perform the…
Juhyeon Park, Heoncheol Lee, Hyuck-Hoon Kwon, Yeji Hwang + 2 more
'Wonseok Choi' 'Wei Yi'] This paper addresses the problem of tracking a high-speed ballistic target in real time. Particle swarm optimization (PSO) can be a solution to overcome the motion of the ballistic target and the nonlinearity of the measurement model. However, in general, particle swarm optimization requires a…
Gangli Liu
problem in an undirected dense graph Authors: ['Gangli Liu'] We provide an efficient ( 2 ) implementation for solving the all pairs minimax path problem or widest path problem in an undirected dense graph. It is a code implementation of the Algorithm 4 (MMJ distance by Calculation and Copy) in a previous paper. The…
Pier Paolo Poier, Louis Lagardère, Jean-Philip Piquemal
We propose a new strategy to solve the Tkatchenko-Scheffler Many-Body Dispersion (MBD) model’s equations. Our approach overcomes the original O(N**3) computational complexity that limits its applicability to large molecular systems within thecontext of O(N) Density Functional Theory (DFT). First, in order to generate…
Pier Paolo Poir, Louis Lagardère, Jean-Philip Piquemal
We propose a new strategy to solve the Tkatchenko-Scheffler Many-Body Dispersion (MBD) model’s equations. Our approach overcomes the original O(N**3) computational complexity that limits its applicability to large molecular systems within thecontext of O(N) Density Functional Theory (DFT). First, in order to generate…
Christopher Myers, Ken Miyazaki, Thomas Trepl, Christine Isborn + 1 more
GPU-accelerated on-the-fly nonadiabatic dynamics is enabled by interfacing the linearized semiclassical dynamics approach with the TeraChem electronic structure program. We describe the computational workflow of the "PySCES" code interface, a Python code for semiclassical dynamics with on-the-fly electronic structure…
Madushanka Manathunga, Hasan Metin Aktulga, Andreas W. Goetz, Kenneth M. Merz + 1 more
We have ported and optimized the GPU accelerated QUICK and AMBER based ab initio QM/MM implementation on AMD GPUs. This encompasses the entire Fock matrix build and force calculation in QUICK including one-electron integrals, two-electron repulsion integrals, exchange-correlation quadrature, and linear algebra…