21 papers · ranked by Valyu relevance
Tim G. J. Rudner, Cong Lu, Michael A. Osborne, Yarin Gal + 1 more
'Yee Whye Teh'] KL-regularized reinforcement learning from expert demonstrations has proved successful in improving the sample efficiency of deep reinforcement learning algorithms, allowing them to be applied to challenging physical real-world tasks. However, we show that KL-regularized reinforcement learning with…
Yifan Zhang, Yifeng Liu, Huizhuo Yuan, Yang Yuan + 2 more
'Andrew C Yao'] Policy gradient algorithms have been successfully applied to enhance the reasoning capabilities of large language models (LLMs). Despite the widespread use of Kullback-Leibler (KL) regularization in policy gradient algorithms to stabilize training, the systematic exploration of how different KL…
Heyang Zhao, Chenlu Ye, Quanquan Gu, Tong Zhang
Reverse-Kullback-Leibler (KL) regularization has emerged to be a predominant technique used to enhance policy optimization in reinforcement learning (RL) and reinforcement learning from human feedback (RLHF), which forces the learned policy to stay close to a reference policy. While the effectiveness and necessity of…
Qingyue Zhao, Kaixuan Ji, Heyang Zhao, Quanquan Gu
\emph{Kullback-Leibler} (KL) regularization is ubiquitous in reinforcement learning algorithms in the form of \emph{reverse} or \emph{forward} KL. Recent studies have demonstrated $ε^{-1}$-type fast rates for decision making under reverse KL regularization, in contrast to the standard $ε^{-2}$-type sample complexity.…
Lingwei Zhu, Zheng Chen, Takamitsu Matsubara, Martha White
Many policy optimization approaches in reinforcement learning incorporate a Kullback-Leilbler (KL) divergence to the previous policy, to prevent the policy from changing too quickly. This idea was initially proposed in a seminal paper on Conservative Policy Iteration, with approximations given by algorithms like TRPO…
John Knight
Latent Factor Analysis via Dynamical Systems (LFADS) is a powerful variational autoencoder for inferring neural population dynamics from spike train data. However, LFADS suffers from pos-terior collapse, where the learned posterior collapses to the prior, eliminating meaningful latent representations. Current solutions…
Constantin Octavian Puiu
In optimization for Machine learning (ML), it is typical that curvature-matrix (CM) estimates rely on an exponential average (EA) of local estimates (giving ea-cm algorithms). This approach has little principled justification, but is very often used in practice. In this paper, we draw a connection between ea-cm…
Matthew Dixon, Tyler Ward, Lizhong Zheng, Chao Tian
Modern computational models in supervised machine learning are often highly parameterized universal approximators. As such, the value of the parameters is unimportant, and only the out of sample performance is considered. On the other hand much of the literature on model estimation assumes that the parameters…
Jürgen Köfinger, Gerhard Hummer
The proper balancing of information from experiment and theory is a long-standing problem in the analysis of noisy and incomplete data. Viewed as a Pareto optimization problem, improved agreement with the experimental data comes at the expense of growing inconsistencies with the theoretical reference model. Here, we…
Jürgen Köfinger, Gerhard Hummer
The proper balancing of information from experiment and theory is a long-standing problem in the analysis of noisy and incomplete data. Viewed as a Pareto optimization problem, improved agreement with the experimental data comes at the expense of growing inconsistencies with the theoretical reference model. Here, we…
Keisuke Ozawa
Statistically weighted principal component analysis (wPCA) is widely used to reduce the noise of scanning transmission electron microscopy-energy-dispersive X-ray (STEM-EDX) spectroscopy data. It is beneficial to retain the spatial resolution of observation in each step of the analysis, but the direct application of…
Jiaji Qiu, Huiying Xu, Xinzhong Zhu, Michael Adjeisah
Multikernel clustering achieves clustering of linearly inseparable data by applying a kernel method to samples in multiple views. A localized SimpleMKKM (LI-SimpleMKKM) algorithm has recently been proposed to perform min-max optimization in multikernel clustering where each instance is only required to be aligned with…
Authors not listed
Electrochemical impedance spectroscopy (EIS) coupled with distribution of relaxation times (DRT) analysis is a robust framework for characterizing electrochemical systems. However, DRT deconvolution is often plagued by spurious peaks, hindering accurate process identification and quantitative parameter estimation. To…
Sarah Friedrich, Andreas Groll, Katja Ickstadt, Thomas Kneib + 3 more
methods and their applications Authors: ['Sarah Friedrich' 'Andreas Groll' 'Katja Ickstadt' 'Thomas Kneib' 'Markus Pauly' 'Jörg Rahnenführer' 'Tim Friede'] A range of regularization approaches have been proposed in the data sciences to overcome overfitting, to exploit sparsity or to improve prediction. Using a broad…
Griffin S. Hampton, Ryan Neff, Zezheng Song, Mustapha Bouhrara + 2 more
Myelin water fraction (MWF) mapping in the central nervous system is a topic of intense research activity. One framework for this requires parameter estimation from a decaying biexponential signal. However, this is often an ill-posed nonlinear problem resulting in unreliable parameter estimates. For linear…
Adrian L. Hauber, Marcus Rosenblatt, Jens Timmer
Ordinary differential equations are frequently employed for mathematical modeling of biological systems. The identification of mechanisms that are specific to certain cell types is crucial for building useful models and to gain insights into the underlying biological processes. Regularization techniques have been…
Denis Tikhonov
Here, we present a new approach for obtaining radial distribution functions (RDF) from the electron diffraction data using a regularized weighted sine least-squares spectral analysis (rwsLSSA). It allows for explicitly transferring the measured experimental uncertainties in the reduced molecular scattering function to…
Antony Mizzi, David M. Walker, Michael Small, José F. F. Mendes
We derive a penalty strength criterion for ridge regression using stochastic complexity, which is a refined variant of the minimum description length principle. Since stochastic complexity does not typically account for the effect of regularization on complexity, despite its ability to simplify models, we are required…
Nurdan Ayse Saran, Fatih Nar, Charles Elkan
This study presents a novel numerical approach that improves the training efficiency of binary logistic regression, a popular statistical model in the machine learning community. Our method achieves training times an order of magnitude faster than traditional logistic regression by employing a novel Soft-Plus…
Radu-Andrei Otopeleanu, Constantin Paleologu, Jacob Benesty, Laura-Maria Dogariu + 3 more
'Laura-Maria Dogariu' 'Cristian-Lucian Stanciu' 'Silviu Ciochină' 'Ka-Fai Cedric Yiu'] The recursive least-squares (RLS) algorithm stands out as an appealing choice in adaptive filtering applications related to system identification problems. This algorithm is able to provide a fast convergence rate for various types…
Søren A. Fuglsang, Kristoffer H. Madsen, Oula Puonti, Hartwig R. Siebner + 1 more
Regression is a principal tool for relating brain responses to stimuli or tasks in computational neuroscience. This often involves fitting linear models with predictors that can be divided into groups, such as distinct stimulus feature subsets in encoding models or features of different neural response channels in…