25 papers · ranked by Valyu relevance
Chaoqiong Fan, Li Yao, Jiacai Zhang, Zonglei Zhen + 1 more
In recent years, brain science and neuroscience have greatly propelled the innovation of computer science. In particular, knowledge from the neurobiology and neuropsychology of the brain revolutionized the development of reinforcement learning (RL) by providing novel interpretable mechanisms of how the brain achieves…
Dong Han, Beni Mulyana, Vladimir Stankovic, Samuel Cheng + 1 more
'Alberto Borboni'] Robotic manipulation challenges, such as grasping and object manipulation, have been tackled successfully with the help of deep reinforcement learning systems. We give an overview of the recent advances in deep reinforcement learning algorithms for robotic manipulation tasks in this review. We begin…
Andreas Nordland, Klaus K. Holst
The R package polle is a unifying framework for learning and evaluating finite stage policies based on observational data. The package implements a collection of existing and novel methods for causal policy learning including doubly robust restricted Q-learning, policy tree learning, and outcome weighted learning. The…
Veronica Chelu, Tom Zahavy, Arthur Guez, Doina Precup + 1 more
'Sebastian Flennerhag'] We work towards a unifying paradigm for accelerating policy optimization methods in reinforcement learning (RL) by integrating foresight in the policy improvement step via optimistic and adaptive updates. Leveraging the connection between policy iteration and policy gradient methods, we view…
Jonah W. Brenner, Chenguang Li, Gabriel Kreiman
Nervous systems learn representations of the world and policies to act within it. We present a framework that uses reward-dependent noise to facilitate policy opti- mization in representation learning networks. These networks balance extracting normative features and task-relevant information to solve tasks. Moreover…
Lucy Lai, Samuel J. Gershman, Thomas Serre
Policy compression is a computational framework that describes how capacity-limited agents trade reward for simpler action policies to reduce cognitive cost. In this study, we present behavioral evidence that humans prefer simpler policies, as predicted by a capacity-limited reinforcement learning model. Across a set…
Adrien Bolland, Gilles Louppe, Damien Ernst
Direct policy optimization in reinforcement learning is usually solved with policy-gradient algorithms, which optimize policy parameters via stochastic gradient ascent. This paper provides a new theoretical interpretation and justification of these algorithms. First, we formulate direct policy optimization in the…
Maryam Sabzevari, Sandor Szedmak, Merja Penttilä, Paula Jouhten + 1 more
Engineered microbial cells present a sustainable alternative to fossil-based synthesis of chemicals and fuels. Cellular synthesis routes are readily assembled and introduced into microbial strains using state-of-the-art synthetic biology tools. However, the optimization of the strains required to reach industrially…
Mingfei Sun, Benjamin J. Ellis, Anuj Mahajan, Sam Devlin + 2 more
Trust Region Policy Optimization (TRPO) is an iterative method that simultaneously maximizes a surrogate objective and enforces a trust region constraint over consecutive policies in each iteration. The combination of the surrogate objective maximization and the trust region enforcement has been shown to be crucial to…
Jonah W. Brenner, Chenguang Li, Gabriel Kreiman
Biological nervous systems learn both internal representations of the world and behavioral policies for acting within it. Motivated by growing evidence that representation learning is a fundamental principle underlying synaptic plasticity, we introduce Neural Stochastic Modulation (NSM): a theory of learning in which…
Weimin Chen, Kelvin Kian Loong Wong, Sifan Long, Zhili Sun + 1 more
'Boris Ryabko'] In the field of reinforcement learning, we propose a Correct Proximal Policy Optimization (CPPO) algorithm based on the modified penalty factor β and relative entropy in order to solve the robustness and stationarity of traditional algorithms. Firstly, In the process of reinforcement learning, this…
Shuze Liu, Samuel Joseph Gershman, Alex Leonidas Doumas
Real-world decision-making often involves navigating large action spaces with state-dependent action values, taxing the limited cognitive resources at our disposal. While previous studies have explored cognitive constraints on generating action consideration sets or refining state-action mappings (policy complexity)…
Matteo Papini, Giorgio Manganini, Alberto Maria Metelli, Marcello Restelli
'Marcello Restelli'] Importance sampling (IS) represents a fundamental technique for a large surge of off-policy reinforcement learning approaches. Policy gradient (PG) methods, in particular, significantly benefit from IS, enabling the effective reuse of previously collected samples, thus increasing sample efficiency.…
Johannes Niediek, Maciej M. Jankowski, Ana Polterovich, Alexander Kazakov + 1 more
Animals can learn complex behaviors. Animal behavior in the lab has traditionally been studied via summary statistics such as trial-based success rates. However, animal behavior is much more fine-grained: a trial in an experiment often consists of multiple actions, and more than one strategy can lead to a successful…
Jianming Wang, Chenyang Lv, Jiting Yin, Yuling Chen + 2 more
Introduction Secondary forests are often characterized by irrational stand structures and spatial imbalances, which constrain the full realization of their ecological functions. Existing deep learning approaches for stand structure optimization still face several limitations, including low computational efficiency…
Shuze Liu, Atsushi Kikumoto, David Badre, Samuel J. Gershman
Making context-dependent decisions incurs cognitive costs. Cognitive control studies have investigated the nature of such costs from both computational and neural perspectives. In this paper, we offer an information-theoretic account of the costs associated with context-dependent decisions. According to this account…
Patrick Sweeney, Jaime Ruiz-Serra, Michael S. Harré, Geert Verdoolaege
A central challenge in artificial intelligence and cognitive science is identifying a unifying principle that governs inference, learning, and action. Active inference proposes such a principle: the minimization of variational free energy. Advocates of active inference argue that the framework subsumes classical models…
Elena Zamaraeva, Christopher M. Collins, Dmytro Antypov, Vladimir V. Gusev + 6 more
Crystal Structure Prediction (CSP) is a fundamental computational problem in materials science. Basin-hopping is a prominent CSP method that combines global Monte Carlo sampling to search over candidate trial structures with local energy minimisation of these candidates. The sampling uses a stochastic policy to…
Authors not listed
Inverse molecular design aims to generate novel chemical structures that satisfy multiple property constraints, yet reinforcement-learning (RL) fine-tuning can be sensitive to how objectives are converted into a scalar reward. Here, we systematically analyze how scalarization choices and stabilization mechanisms shape…
Octave Oliviers, Glenn Vinnicombe
The asymptotic behaviour of Monte Carlo optimistic policy iteration (MC-O-PI) is a long-standing open question. When the model of the environment is unknown, as is common in practice, the only known condition that guarantees convergence to optimality is impractical. In its canonical form, this condition requires that…
Russell Jeter, Dmitrii Todorov, Yaroslav Molkov
A clinician guiding a stroke patient through a 45-minute rehabilitation session, a coach planning a training day, a teacher choosing the order of practice problems, they all face the same question: “given everything practiced so far, what should the next trial be?” The motor-learning literature offers two coarse…
Giorgio Taricco, Miguel Rubi
Entropy regularization is a recurring mechanism in reinforcement learning (RL), but its meaning changes across algorithmic settings. In classical online RL, entropy encourages exploration and smooths policy improvement; in inverse RL and imitation learning, maximum-entropy resolves ambiguity among expert-consistent…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Benson Chen, Xiang Fu, Tommi Jaakkola, Regina Barzilay
Searching for novel molecular compounds with desired properties is an important problem in drug discovery. Many existing frameworks generate molecules one atom at a time. We instead propose a flexible editing paradigm that generates molecules using learned molecular fragments---meaningful substructures of molecules. To…
Jeff Guo, Philippe Schwaller
Sample efficiency is a fundamental challenge in de novo molecular design. Ideally, molecular generative models should learn to satisfy desired objectives under minimal oracle evaluations (computational prediction or wet-lab experiment). This problem becomes more apparent when using oracles that can provide increased…