Search · four archives
Search · four archives
24 papers · ranked by Valyu relevance
Benyamin Ghojogh, Ali Ghodsi, Fakhri Karray, Mark Crowley
This is a tutorial and survey paper on Karush-Kuhn-Tucker (KKT) conditions, first-order and second-order numerical optimization, and distributed optimization. After a brief review of history of optimization, we start with some preliminaries on properties of sets, norms, functions, and concepts of optimization. Then, we…
Aline C. Soterroni, Roberto L. Galski, Marluce C. Scarabello, Fernando M. Ramos
'Fernando M. Ramos'] In this work, the q-Gradient (q-G) method, a q-version of the Steepest Descent method, is presented. The main idea behind the q-G method is the use of the negative of the q-gradient vector of the objective function as the search direction. The q-gradient vector, or simply the q-gradient, is a…
Alexandra Zverovich, Matthew Hutchings, Bertrand Gauthier
We investigate the properties of a class of piecewise-fractional maps arising from the introduction of an invariance under rescaling into convex quadratic maps. The subsequent maps are quasiconvex, and pseudoconvex on specific convex cones; they can be optimised via exact line search along admissible directions. We…
Esmail Abdul Fattah, Janet van Niekerk, Håvard Rue
Computing the gradient of a function provides fundamental information about its behavior. This information is essential for several applications and algorithms across various fields. One common application that require gradients are optimization techniques such as stochastic gradient descent, Newton's method and trust…
MOHAMED HASSAN, ALEKSANDAR VAKANSKI, BOYU ZHANG, MIN XIAN
The generalization performance of deep neural networks (DNNs) is a critical factor in achieving robust model behavior on unseen data. Recent studies have highlighted the importance of sharpness-based measures in promoting generalization by encouraging convergence to flatter minima. Among these approaches…
Lulu He, Yanan Du, Jianchao Bai
The conjugate gradient method is widely recognized as a foundational technique for large-scale unconstrained optimization. In this work, we introduce an Accelerated Stochastic Conjugate Gradient (ASCG) algorithm, specifically designed for a class of convex empirical risk minimization problems. The proposed ASCG method…
Leonard Schmiester, Daniel Weindl, Jan Hasenauer
Unknown parameters of dynamical models are commonly estimated from experimental data. However, while various efficient optimization and uncertainty analysis methods have been proposed for quantitative data, methods for qualitative data are rare and suffer from bad scaling and convergence. Here, we propose an efficient…
Simone Carlo Surace, Jean-Pascal Pfister, Wulfram Gerstner, Johanni Brea + 1 more
'Johanni Brea' 'Francis Ouellette'] This is a PLOS Computational Biology Education paper. The idea that the brain functions so as to minimize certain costs pervades theoretical neuroscience. Because a cost function by itself does not predict how the brain finds its minima, additional assumptions about the optimization…
Ming Tian, Hui-Fang Zhang
Let H be a real Hilbert space and C be a nonempty closed convex subset of H. Assume that g is a real-valued convex function and the gradient ∇g is \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs}…
Akhil Shajan, Madushanka Manathunga, Andreas Goetz, Kenneth Merz
Based on a series of energy minimizations with starting structures obtained from the Baker test set of 30 organic molecules, a comparison is made between various open- source geometry optimization codes that are interfaced with the open-source QUantum Interaction Computational Kernel (QUICK) program for gradient and…
Jiawei Zhang
In this paper, we aim at providing an introduction to the gradient descent based optimization algorithms for learning deep neural network models. Deep learning models involving multiple nonlinear projection layers are very challenging to train. Nowadays, most of the deep learning model training still relies on the back…
AKHIL SHAJAN, Madushanka Manathunga, Andreas Goetz, Kenneth Merz
Based on a series of energy minimizations with starting structures obtained from the Baker test set of 30 organic molecules, a comparison is made between various open-source geometry optimization codes that are interfaced with the open-source QUantum Interaction Computational Kernel (QUICK) program for gradient and…
Kazunori D Yamada
In the deep learning era, a gradient descent method is the most common method to optimize parameters of neural networks. Among various mathematical optimization methods, a gradient descent method is the most naive method. Although controlling a learning rate of the method is necessary for quick convergence, the…
Hugo Silva, Martha White
Network? Authors: ['Hugo Silva' 'Martha White'] Oftentimes, machine learning applications using neural networks involve solving discrete optimization problems, such as in pruning, parameter-isolation-based continual learning and training of binary networks. Still, these discrete problems are combinatorial in nature and…
Akhil Shajan, Madushanka Manathunga, Andreas Goetz, Kenneth Merz
Based on a series of energy minimizations with starting structures obtained from the Baker test set of 30 organic molecules, a comparison is made between various open source geometry optimization codes that are interfaced with the open-source QUantum Interaction Computational Kernel (QUICK) program for gradient and…
Aline C. Soterroni, Roberto Luiz Galski, Fernando M. Ramos
The q-gradient is an extension of the classical gradient vector based on the concept of Jackson's derivative. Here we introduce a preliminary version of the q-gradient method for unconstrained global optimization. The main idea behind our approach is the use of the negative of the q-gradient of the objective function…
Aaron Defazio, Baoyu Zhou, Lin Xiao
The classical AdaGrad method adapts the learning rate by dividing by the square root of a sum of squared gradients. Because this sum on the denominator is increasing, the method can only decrease step sizes over time, and requires a learning rate scaling hyper-parameter to be carefully tuned. To overcome this…
Nicolas García Trillos, Zachary Kaplan, Daniel Sanz-Alonso
The aim of this paper is to provide new theoretical and computational understanding on two loss regularizations employed in deep learning, known as local entropy and heat regularization. For both regularized losses, we introduce variational characterizations that naturally suggest a two-step scheme for their…
Georg Hahn, Sharon M. Lutz, Nilanjana Laha, Michael Cho + 2 more
High dimensional linear regression problems are often fitted using LASSO-type approaches. Although the LASSO objective function is convex, it is not differentiable everywhere, making the use of gradient descent methods for minimization not straightforward. To avoid this technical issue, we apply Nesterov smoothing to…
Abbas Kazemipour, Behtash Babadi, Min Wu, Kaspar Podgorski + 1 more
We consider the problem of optimizing general convex objective functions with nonnegativity constraints. Using the Karush-Kuhn-Tucker (KKT) conditions for the nonnegativity constraints we will derive fast multiplicative update rules for several problems of interest in signal processing, including non-negative…
Eric Hermes, Khachik Sargsyan, Habib Najm, Judit Zádor
We present a new algorithm for the optimization of molecular structures to saddle points on the potential energy surface using a redundant internal coordinate system. This algorithm automates the procedure of defining the internal coordinate system, including the handling of linear bending angles, e.g. through the…
Georg Hahn, Sharon M. Lutz, Nilanjana Laha, Christoph Lange
Penalized linear regression approaches that include an L_1_ term have become an important tool in statistical data analysis. One prominent example is the least absolute shrinkage and selection operator (Lasso), though the class of L_1_ penalized regression operators also includes the fused and graphical Lasso, the…
Lucian Chan, Geoffrey Hutchison, Garrett Morris
Generating low-energy molecular conformers is a key task for many areas of computational chemistry, molecular modeling and cheminformatics. Most current conformer generation methods primarily focus on generating geometrically diverse conformers rather than finding the most probable or energetically lowest minima. Here…
Authors not listed
Solving optimization problems, especially for nonlinear and constrained systems, is a challenge. Decades of specialized algorithms have been developed for general and special cases of root finding, minimization (including constraints), for parameter estimation, and mapping connected spaces. These approaches typically…