23 papers · ranked by Valyu relevance
Chaoqiong Fan, Li Yao, Jiacai Zhang, Zonglei Zhen + 1 more
In recent years, brain science and neuroscience have greatly propelled the innovation of computer science. In particular, knowledge from the neurobiology and neuropsychology of the brain revolutionized the development of reinforcement learning (RL) by providing novel interpretable mechanisms of how the brain achieves…
Veronica Chelu, Tom Zahavy, Arthur Guez, Doina Precup + 1 more
'Sebastian Flennerhag'] We work towards a unifying paradigm for accelerating policy optimization methods in reinforcement learning (RL) by integrating foresight in the policy improvement step via optimistic and adaptive updates. Leveraging the connection between policy iteration and policy gradient methods, we view…
Jonah W. Brenner, Chenguang Li, Gabriel Kreiman
Nervous systems learn representations of the world and policies to act within it. We present a framework that uses reward-dependent noise to facilitate policy opti- mization in representation learning networks. These networks balance extracting normative features and task-relevant information to solve tasks. Moreover…
Adrien Bolland, Gilles Louppe, Damien Ernst
Direct policy optimization in reinforcement learning is usually solved with policy-gradient algorithms, which optimize policy parameters via stochastic gradient ascent. This paper provides a new theoretical interpretation and justification of these algorithms. First, we formulate direct policy optimization in the…
Mingfei Sun, Benjamin J. Ellis, Anuj Mahajan, Sam Devlin + 2 more
Trust Region Policy Optimization (TRPO) is an iterative method that simultaneously maximizes a surrogate objective and enforces a trust region constraint over consecutive policies in each iteration. The combination of the surrogate objective maximization and the trust region enforcement has been shown to be crucial to…
Feng Zhang, Jiang Li, Ye Wang, Lihong Guo + 4 more
'Hongwei Zhao' 'Carlo Alberto Avizzano'] Capability assessment plays a crucial role in the demonstration and construction of equipment. To improve the accuracy and stability of capability assessment, we study the neural network learning algorithms in the field of capability assessment and index sensitivity. Aiming at…
Samuel J. Gershman
When humans and other animals make repeated choices, they tend to repeat previously chosen actions independently of their reward history. This paper locates the origin of perseveration in a trade-off between two computational goals: maximizing rewards and minimizing the complexity of the action policy. We develop an…
Jonah W. Brenner, Chenguang Li, Gabriel Kreiman
Biological nervous systems learn both internal representations of the world and behavioral policies for acting within it. Motivated by growing evidence that representation learning is a fundamental principle underlying synaptic plasticity, we introduce Neural Stochastic Modulation (NSM): a theory of learning in which…
Boris Belousov, Jan Peters
An optimal feedback controller for a given Markov decision process (MDP) can in principle be synthesized by value or policy iteration. However, if the system dynamics and the reward function are unknown, a learning agent must discover an optimal controller via direct interaction with the environment. Such interactive…
Samuel J. Gershman, Lucy Lai
Action selection requires a policy that maps states of the world to a distribution over actions. The amount of memory needed to specify the policy (the policy complexity) increases with the state-dependence of the policy. If there is a capacity limit for policy complexity, then there will also be a trade-off between…
Shijun Wang, Baocheng Zhu, Chen Li, Mingzhe Wu + 3 more
'Wei Chu' 'Qi Yuan'] In this paper, We propose a general Riemannian proximal optimization algorithm with guaranteed convergence to solve Markov decision process (MDP) problems. To model policy functions in MDP, we employ Gaussian mixture model (GMM) and formulate it as a nonconvex optimization problem in the Riemannian…
Jeffrey F. Queißer, Jochen J. Steil
Modern robotic applications create high demands on adaptation of actions with respect to variance in a given task. Reinforcement learning is able to optimize for these changing conditions, but relearning from scratch is hardly feasible due to the high number of required rollouts. We propose a parameterized skill that…
Yonathan Efroni, Lior Shani, Aviv Rosenberg, Shie Mannor
Policy optimization methods are one of the most widely used classes of Reinforcement Learning (RL) algorithms. Yet, so far, such methods have been mostly analyzed from an optimization perspective, without addressing the problem of exploration, or by making strong assumptions on the interaction with the environment. In…
Andreas Nordland, Klaus K. Holst
The R package polle is a unifying framework for learning and evaluating finite stage policies based on observational data. The package implements a collection of existing and novel methods for causal policy learning including doubly robust restricted Q-learning, policy tree learning, and outcome weighted learning. The…
Weimin Chen, Kelvin Kian Loong Wong, Sifan Long, Zhili Sun + 1 more
'Boris Ryabko'] In the field of reinforcement learning, we propose a Correct Proximal Policy Optimization (CPPO) algorithm based on the modified penalty factor β and relative entropy in order to solve the robustness and stationarity of traditional algorithms. Firstly, In the process of reinforcement learning, this…
Elena Zamaraeva, Christopher M. Collins, Dmytro Antypov, Vladimir V. Gusev + 6 more
Crystal Structure Prediction (CSP) is a fundamental computational problem in materials science. Basin-hopping is a prominent CSP method that combines global Monte Carlo sampling to search over candidate trial structures with local energy minimisation of these candidates. The sampling uses a stochastic policy to…
Authors not listed
Inverse molecular design aims to generate novel chemical structures that satisfy multiple property constraints, yet reinforcement-learning (RL) fine-tuning can be sensitive to how objectives are converted into a scalar reward. Here, we systematically analyze how scalarization choices and stabilization mechanisms shape…
Octave Oliviers, Glenn Vinnicombe
The asymptotic behaviour of Monte Carlo optimistic policy iteration (MC-O-PI) is a long-standing open question. When the model of the environment is unknown, as is common in practice, the only known condition that guarantees convergence to optimality is impractical. In its canonical form, this condition requires that…
Russell Jeter, Dmitrii Todorov, Yaroslav Molkov
A clinician guiding a stroke patient through a 45-minute rehabilitation session, a coach planning a training day, a teacher choosing the order of practice problems, they all face the same question: “given everything practiced so far, what should the next trial be?” The motor-learning literature offers two coarse…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Jeff Guo, Philippe Schwaller
Sample efficiency is a fundamental challenge in de novo molecular design. Ideally, molecular generative models should learn to satisfy desired objectives under minimal oracle evaluations (computational prediction or wet-lab experiment). This problem becomes more apparent when using oracles that can provide increased…
Tom Lefebvre, Guillaume Crevecoeur
In this article, we present a generalized view on Path Integral Control (PIC) methods. PIC refers to a particular class of policy search methods that are closely tied to the setting of Linearly Solvable Optimal Control (LSOC), a restricted subclass of nonlinear Stochastic Optimal Control (SOC) problems. This class is…
Benson Chen, Xiang Fu, Tommi Jaakkola, Regina Barzilay
Searching for novel molecular compounds with desired properties is an important problem in drug discovery. Many existing frameworks generate molecules one atom at a time. We instead propose a flexible editing paradigm that generates molecules using learned molecular fragments---meaningful substructures of molecules. To…