14 papers · ranked by Valyu relevance
Chaoqiong Fan, Li Yao, Jiacai Zhang, Zonglei Zhen + 1 more
In recent years, brain science and neuroscience have greatly propelled the innovation of computer science. In particular, knowledge from the neurobiology and neuropsychology of the brain revolutionized the development of reinforcement learning (RL) by providing novel interpretable mechanisms of how the brain achieves…
Lucy Lai, Samuel J. Gershman, Thomas Serre
Policy compression is a computational framework that describes how capacity-limited agents trade reward for simpler action policies to reduce cognitive cost. In this study, we present behavioral evidence that humans prefer simpler policies, as predicted by a capacity-limited reinforcement learning model. Across a set…
Feng Zhang, Jiang Li, Ye Wang, Lihong Guo + 4 more
'Hongwei Zhao' 'Carlo Alberto Avizzano'] Capability assessment plays a crucial role in the demonstration and construction of equipment. To improve the accuracy and stability of capability assessment, we study the neural network learning algorithms in the field of capability assessment and index sensitivity. Aiming at…
Gerhard Neumann, Christian Daniel, Alexandros Paraschos, Andras Kupcsik + 1 more
'Andras Kupcsik' 'Jan Peters'] A promising idea for scaling robot learning to more complex tasks is to use elemental behaviors as building blocks to compose more complex behavior. Ideally, such building blocks are used in combination with a learning algorithm that is able to learn to select, adapt, sequence and…
Shuze Liu, Samuel Joseph Gershman, Alex Leonidas Doumas
Real-world decision-making often involves navigating large action spaces with state-dependent action values, taxing the limited cognitive resources at our disposal. While previous studies have explored cognitive constraints on generating action consideration sets or refining state-action mappings (policy complexity)…
Boris Belousov, Jan Peters
An optimal feedback controller for a given Markov decision process (MDP) can in principle be synthesized by value or policy iteration. However, if the system dynamics and the reward function are unknown, a learning agent must discover an optimal controller via direct interaction with the environment. Such interactive…
Jeffrey F. Queißer, Jochen J. Steil
Modern robotic applications create high demands on adaptation of actions with respect to variance in a given task. Reinforcement learning is able to optimize for these changing conditions, but relearning from scratch is hardly feasible due to the high number of required rollouts. We propose a parameterized skill that…
Jianming Wang, Chenyang Lv, Jiting Yin, Yuling Chen + 2 more
Introduction Secondary forests are often characterized by irrational stand structures and spatial imbalances, which constrain the full realization of their ecological functions. Existing deep learning approaches for stand structure optimization still face several limitations, including low computational efficiency…
Weimin Chen, Kelvin Kian Loong Wong, Sifan Long, Zhili Sun + 1 more
'Boris Ryabko'] In the field of reinforcement learning, we propose a Correct Proximal Policy Optimization (CPPO) algorithm based on the modified penalty factor β and relative entropy in order to solve the robustness and stationarity of traditional algorithms. Firstly, In the process of reinforcement learning, this…
Ronald Ortner
We consider a reinforcement learning setting where the learner is given a set of possible models containing the true model. While there are algorithms that are able to successfully learn optimal behavior in this setting, they do so without trying to identify the underlying true model. Indeed, we show that there are…
Samuel J. Gershman, Lucy Lai
Action selection requires a policy that maps states of the world to a distribution over actions. The amount of memory needed to specify the policy (the policy complexity) increases with the state-dependence of the policy. If there is a capacity limit for policy complexity, then there will also be a trade-off between…
Giorgio Taricco, Miguel Rubi
Entropy regularization is a recurring mechanism in reinforcement learning (RL), but its meaning changes across algorithmic settings. In classical online RL, entropy encourages exploration and smooths policy improvement; in inverse RL and imitation learning, maximum-entropy resolves ambiguity among expert-consistent…
Patrick Sweeney, Jaime Ruiz-Serra, Michael S. Harré, Geert Verdoolaege
A central challenge in artificial intelligence and cognitive science is identifying a unifying principle that governs inference, learning, and action. Active inference proposes such a principle: the minimization of variational free energy. Advocates of active inference argue that the framework subsumes classical models…
Tom Lefebvre, Guillaume Crevecoeur
In this article, we present a generalized view on Path Integral Control (PIC) methods. PIC refers to a particular class of policy search methods that are closely tied to the setting of Linearly Solvable Optimal Control (LSOC), a restricted subclass of nonlinear Stochastic Optimal Control (SOC) problems. This class is…