13 papers · ranked by Valyu relevance
Sebastian Ruder
Gradient descent optimization algorithms, while increasingly popular, are often used as black-box optimizers, as practical explanations of their strengths and weaknesses are hard to come by. This article aims to provide the reader with intuitions with regard to the behaviour of different algorithms that will allow her…
Benyamin Ghojogh, Ali Ghodsi, Fakhri Karray, Mark Crowley
This is a tutorial and survey paper on Karush-Kuhn-Tucker (KKT) conditions, first-order and second-order numerical optimization, and distributed optimization. After a brief review of history of optimization, we start with some preliminaries on properties of sets, norms, functions, and concepts of optimization. Then, we…
Esmail Abdul Fattah, Janet van Niekerk, Håvard Rue
Computing the gradient of a function provides fundamental information about its behavior. This information is essential for several applications and algorithms across various fields. One common application that require gradients are optimization techniques such as stochastic gradient descent, Newton's method and trust…
Shiliang Sun, Zehui Cao, Zhu Han, Jing Zhao
—Machine learning develops rapidly, which has made many theoretical breakthroughs and is widely applied in various fields. Optimization, as an important part of machine learning, has attracted much attention of researchers. With the exponential growth of data amount and the increase of model complexity, optimization…
Pushparaja Murugan, Shanmugasundaram Durairaj
Convolution Neural Networks, known as ConvNets exceptionally perform well in many complex machine learning tasks. The architecture of ConvNets demands the huge and rich amount of data and involves with a vast number of parameters that leads the learning takes to be computationally expensive, slow convergence towards…
Saeed Asadi, Sonia Gharibzadeh, Shiva Zangeneh, Masoud Reihanifar + 2 more
Multidimensional Surface 3D Visualizations and Initial Point Sensitivity Authors: ['Saeed Asadi' 'Sonia Gharibzadeh' 'Shiva Zangeneh' 'Masoud Reihanifar' 'Mehrzad Rahimi' 'Lazim Abdullah'] This study examines several renowned gradient-based optimization techniques and focuses on their computational efficiency and…
Jian Wu, Matthias Poloczek, Andrew Gordon Wilson, Peter I. Frazier
Bayesian optimization has been successful at global optimization of expensiveto-evaluate multimodal objective functions. However, unlike most optimization methods, Bayesian optimization typically does not use derivative information. In this paper we show how Bayesian optimization can exploit derivative information to…
Kaustubh Yadav
—One of the most important parts of Artificial Neural Networks is minimizing the loss functions which tells us how good or bad our model is. To minimize these losses we need to tune the weights and biases. Also to calculate the minimum value of a function we need gradient. And to update our weights we need gradient…
Chad Kelterborn, Marcin Mazur, Bogdan V. Petrenko
Gradient descent algorithms have been used in countless applications since the inception of Newton's method. The explosion in the number of applications of neural networks has re-energized efforts in recent years to improve the standard gradient descent method in both efficiency and accuracy. These methods modify the…
Naoki Sato, Hideaki Iiduka
Graduated optimization is a global optimization technique that is used to minimize a multimodal nonconvex function by smoothing the objective function with noise and gradually refining the solution. This paper experimentally evaluates the performance of the explicit graduated optimization algorithm with an optimal…
Dariush Bahrami, Sadegh Pouriyan Zadeh
We introduce Gravity, another algorithm for gradient-based optimization. In this paper, we explain how our novel idea change parameters to reduce the deep learning model's loss. It has three intuitive hyper-parameters that the best values for them are proposed. Also, we propose an alternative to moving average. To…
Simone Carlo Surace, Johanni Brea
The idea that the brain functions so as to minimize certain costs pervades theoretical neuroscience. Since a cost function by itself does not predict how the brain finds its minima, additional assumptions about the optimization method need to be made to predict the dynamics of physiological quantities. In this context…
Ramkrishna Acharya
This study reviews popular stochastic gradient-based schemes based on large least-square problems. These schemes, often called optimizers in machine learning, play a crucial role in finding better model parameters. Hence, this study focuses on viewing such optimizers with different hyper-parameters and analyzing them…