24 papers · ranked by Valyu relevance
Marco Boresta, Tommaso Colombo, Alberto De Santis, Stefano Lucidi
In this paper we focus on the linear functionals defining an approximate version of the gradient of a function. These functionals are often used when dealing with optimization problems where the computation of the gradient of the objective function is costly or the objective function values are affected by some noise.…
I. D. Coope, Rachael Tappenden
Simplex gradients are an essential feature of many derivative free optimization algorithms, and can be employed, for example, as part of the process of defining a direction of search, or as part of a termination criterion. The calculation of a general simplex gradient in Rn can be computationally expensive, and often…
Stephen Jay Gould, Ming Xu, Zhiwei Xu, Yanbin Liu
We explore conditions for when the gradient of a deep declarative node can be approximated by ignoring constraint terms and still result in a descent direction for the global loss function. This has important practical application when training deep learning models since the approximation is often computationally much…
Yuri K. Shestopaloff, Alexander Y. Shestopaloff
The problems of computational data processing involving regression, interpolation, reconstruction and imputation for multidimensional big datasets are becoming more important these days, because of the availability of data and their widely spread usage in business, technological, scientific and other applications. The…
Albert S. Berahas, Liyuan Cao, Krzysztof Choromański, Katya Scheinberg
'Katya Scheinberg'] In this paper, we consider derivative free optimization problems, where the objective function is smooth but is computed with some amount of noise, the function evaluations are expensive and no derivative information is available. We are motivated by policy optimization problems in reinforcement…
I. D. Coope, Rachael Tappenden
This work investigates finite differences and the use of interpolation models to obtain approximations to the first and second derivatives of a function. Here, it is shown that if a particular set of points is used in the interpolation model, then the solution to the associated linear system (i.e., approximations to…
Kai Zhang, Zengfei Wang, Liming Zhang, Jun Yao + 2 more
'Lixiang Li'] In this paper, we investigate the application of a new method, the Finite Difference and Stochastic Gradient (Hybrid method), for history matching in reservoir models. History matching is one of the processes of solving an inverse problem by calibrating reservoir models to dynamic behaviour of the…
Joachim Almquist, Jacob Leander, Mats Jirstrand
The first order conditional estimation (FOCE) method is still one of the parameter estimation workhorses for nonlinear mixed effects (NLME) modeling used in population pharmacokinetics and pharmacodynamics. However, because this method involves two nested levels of optimizations, with respect to the empirical Bayes…
Zelin Pei, Xiaoyu He, Yi Pan, Baichun Peng + 2 more
Black-box stochastic optimization involves sampling in both the solution and data spaces. Traditional variance reduction methods mainly designed for reducing the data sampling noise may suffer from slow convergence if the noise in the solution space is poorly handled. In this paper, we present a novel zeroth-order…
Xiaoqing Huang, Andersen Ang, Aatman Pushkarkumar, Kun Huang + 2 more
Inferring gene regulation from time-course expression profiles is essential for understanding how cells transition between states during development, differentiation, and disease progression. Existing approaches often model expression dynamics with ordinary differential equations (ODEs). However, due to the…
Akhil Shajan, Madushanka Manathunga, Andreas Goetz, Kenneth Merz
Based on a series of energy minimizations with starting structures obtained from the Baker test set of 30 organic molecules, a comparison is made between various open- source geometry optimization codes that are interfaced with the open-source QUantum Interaction Computational Kernel (QUICK) program for gradient and…
Michael Hutcheon, Andrew Teale
Algorithms are presented for performing a topological analysis of an arbitrary function, evaluated on an arbitrary grid of points. These algorithms work strictly by post-processing the data and require no additional function evaluations. This is achieved by connecting the grid points with a neighbourhood graph…
AKHIL SHAJAN, Madushanka Manathunga, Andreas Goetz, Kenneth Merz
Based on a series of energy minimizations with starting structures obtained from the Baker test set of 30 organic molecules, a comparison is made between various open-source geometry optimization codes that are interfaced with the open-source QUantum Interaction Computational Kernel (QUICK) program for gradient and…
A. V. Lobanov
Higher Order Smoothness Function Condition Authors: ['A. V. Lobanov'] Abstract. This paper is devoted to the study (common in many applications) of the black-box optimization problem, where the black-box represents a gradient-free oracle ˜f = f(x) + ξ providing the objective function value with some stochastic noise.…
Akhil Shajan, Madushanka Manathunga, Andreas Goetz, Kenneth Merz
Based on a series of energy minimizations with starting structures obtained from the Baker test set of 30 organic molecules, a comparison is made between various open source geometry optimization codes that are interfaced with the open-source QUantum Interaction Computational Kernel (QUICK) program for gradient and…
Seyedeh Azadeh Fallah Mortezanejad, Ali Mohammad-Djafari, Antonio M. Scarfone, Volker J Schmid + 1 more
'Antonio M. Scarfone' 'Volker J Schmid' 'Zahra Amini Farsani'] In any Bayesian computations, the first step is to derive the joint distribution of all the unknown variables given the observed data. Then, we have to do the computations. There are four general methods for performing computations: Joint MAP optimization…
Aref Miri Rekavandi, Saad Jbabdi, Stephen M. Smith
This paper presents a framework for modelling the topography of whole-brain connectivity in resting-state functional MRI. The aim is to disentangle functional segregation, which manifests as abrupt changes in connectivity, from so-called gradients, i.e., smooth variations in connectivity across the brain. Our core…
Krishna Rijal, Pankaj Mehta
The Gillespie algorithm is commonly used to simulate and analyze complex chemical reaction networks. Here, we leverage recent breakthroughs in deep learning to develop a fully differentiable variant of the Gillespie algorithm. The differentiable Gillespie algorithm (DGA) approximates discontinuous operations in the…
H. van de Beek, M. Beldjenna, M. Fidler, L.B. Zwep + 1 more
Asymptotic standard errors for the parameters of a nonlinear mixed-effects model fitted by first-order conditional estimation (FOCE) or FOCE with interaction (FOCEI) require the observed (Fisher) information — the negative second derivative of the population objective at the optimum. The gradient of this objective can…
Philipp Frank, Reimar Leike, Torsten A. Enßlin, Carlos Alberto De Bragança Pereira
'Carlos Alberto De Bragança Pereira'] Efficiently accessing the information contained in non-linear and high dimensional probability distributions remains a core challenge in modern statistics. Traditionally, estimators that go beyond point estimates are either categorized as Variational Inference (VI) or Markov-Chain…
Kazunori D Yamada
In the deep learning era, a gradient descent method is the most common method to optimize parameters of neural networks. Among various mathematical optimization methods, a gradient descent method is the most naive method. Although controlling a learning rate of the method is necessary for quick convergence, the…
Georg Hahn, Sharon M. Lutz, Nilanjana Laha, Michael Cho + 2 more
High dimensional linear regression problems are often fitted using LASSO-type approaches. Although the LASSO objective function is convex, it is not differentiable everywhere, making the use of gradient descent methods for minimization not straightforward. To avoid this technical issue, we apply Nesterov smoothing to…
Philipp Grohs, Markus Sprecher, Thomas Yu
We consider the problem of approximating a function f from an Euclidean domain to a manifold M by scattered samples $f\xi _i_{i\in \mathcal{I}}$, where the data sites $\xi _i_{i\in \mathcal{I}}$ are assumed to be locally close but can otherwise be far apart points scattered throughout the domain. We introduce a natural…
Rebecca K. Borchering, Scott A. McKinley
In the last decade there has been growing criticism of the use of Stochastic Differential Equations (SDEs) to approximate discrete state-space, continuous-time Markov chain population models. In particular, several authors have demonstrated the failure of Diffusion Approximation, as it is often called, to approximate…