24 papers · ranked by Valyu relevance
Yuezhongyi Sun, Boyu Yang, Mehmet Cunkas
In the dynamic field of deep reinforcement learning, the self-attention mechanism has been increasingly recognized. Nevertheless, its application in discrete problem domains has been relatively limited, presenting complex optimization challenges. This article introduces a pioneering deep reinforcement learning…
Yoshimasa Kubo, Eric Chalmers, Artur Luczak
Backpropagation has been used to train neural networks for many years, allowing them to solve a wide variety of tasks like image classification, speech recognition, and reinforcement learning tasks. But the biological plausibility of backpropagation as a mechanism of neural learning has been questioned. Equilibrium…
Benton Girdler, William Caldbeck, Jihye Bae
Creating flexible and robust brain machine interfaces (BMIs) is currently a popular topic of research that has been explored for decades in medicine, engineering, commercial, and machine-learning communities. In particular, the use of techniques using reinforcement learning (RL) has demonstrated impressive results but…
Zuyue Fu, Zhuoran Yang, Zhaoran Wang
We study the global convergence and global optimality of actor-critic, one of the most popular families of reinforcement learning algorithms. While most existing works on actor-critic employ bi-level or two-timescale updates, we focus on the more practical single-timescale setting, where the actor and critic are…
Shalabh Bhatnagar, Vivek S. Borkar, Soumyajit Guin
We revisit the standard formulation of tabular actor-critic algorithm as a two time-scale stochastic approximation with value function computed on a faster time-scale and policy computed on a slower time-scale. This emulates policy iteration. We observe that reversal of the time scales will in fact emulate value…
Liyuan Zheng, Tanner Fiez, Zane Alumbaugh, Benjamin J. Chasnov + 1 more
'Lillian J. Ratliff'] The hierarchical interaction between the actor and critic in actor-critic based reinforcement learning algorithms naturally lends itself to a game-theoretic interpretation. We adopt this viewpoint and model the actor and critic interaction as a two-player general-sum game with a leader-follower…
Anushka Deshpande
The aim of this paper is twofold. First, it seeks to uncover the algorithms that humans and other animals employ for learning in decision-making strategies within non-zero-sum games, specifically focusing on fully observable iterated prisoner’s dilemma scenarios. Second, it aims to develop a new model to explain…
Simone Parisi, Voot Tangkaratt, Jan Peters, Mohammad Emtiyaz Khan
Actor-critic methods can achieve incredible performance on difficult reinforcement learning problems, but they are also prone to instability. This is partly due to the interaction between the actor and the critic during learning, e.g., an inaccurate step taken by one of them might adversely affect the other and…
Carlos Antolín, Joan Falcó-Roget, Luis Serrano-Fernández, Néstor Parga
Decision-making depends on coordinated computations distributed across dorsal and ventral circuits, often described as actor–critic systems. We examined how this division of labor gives rise to value- and decision-related representations by training a bimodular recurrent network on an economic choice task and…
Junfeng Wen, Saurabh Kumar, Ramki Gummadi, Dale Schuurmans
Actor-critic (AC) methods are ubiquitous in reinforcement learning. Although it is understood that AC methods are closely related to policy gradient (PG), their precise connection has not been fully characterized previously. In this paper, we explain the gap between AC and PG methods by identifying the exact adjustment…
Sukriti Verma, Ayush Chopra, Jayakumar Subramanian, Mausoom Sarkar + 3 more
'Nikaash Puri' 'Piyush Gupta' 'Balaji Krishnamurthy'] The two-time scale nature of SAC, which is an actor-critic algorithm, is characterised by the fact that the critic estimate has not converged for the actor at any given time, but since the critic learns faster than the actor, it ensures eventual consistency between…
Nicolas Frémaux, Henning Sprekeler, Wulfram Gerstner, Lyle J. Graham
Animals repeat rewarded behaviors, but the physiological basis of reward-based learning has only been partially elucidated. On one hand, experimental evidence shows that the neuromodulator dopamine carries information about rewards and affects synaptic plasticity. On the other hand, the theory of reinforcement learning…
Menghao Wu, Yanbin Gao, Alexander Jung, Qiang Zhang + 1 more
Model-free reinforcement learning is a powerful and efficient machine-learning paradigm which has been generally used in the robotic control domain. In the reinforcement learning setting, the value function method learns policies by maximizing the state-action value (Q value), but it suffers from inaccurate Q…
Mo Zhou, Jianfeng Lu
We propose a single timescale actor-critic algorithm to solve the linear quadratic regulator (LQR) problem. A least squares temporal difference (LSTD) method is applied to the critic and a natural policy gradient method is used for the actor. We give a proof of convergence with sample complexity O(ε −1 log(ε −1 ) 2 ).…
Charline Tessereau, Reuben O’Dea, Stephen Coombes, Tobias Bast
Humans and non-human animals show great flexibility in spatial navigation, including the ability to return to specific locations based on as few as one single experience. To study spatial navigation in the laboratory, watermaze tasks, in which rats have to find a hidden platform in a pool of cloudy water surrounded by…
Cesar Guevara, Bilal Alatas
Currently, the stock market is attractive, and it is challenging to develop an efficient investment model with high accuracy due to changes in the values of the shares for political, economic, and social reasons. This article presents an innovative proposal for a short-term, automatic investment model to reduce capital…
Massimo Silvetti, Eliana Vassena, Elger Abrahamse, Tom Verguts
The dorsal anterior cingulate cortex (dACC) is central in higher-order cognition and behavioural flexibility. The computational nature of this region, however, has remained elusive. Here we propose a new model – the Reinforcement Meta Learner (RML) – based on the bidirectional anatomical connections of the ACC with…
Jean-Paul Noel, Ruiyi Zhang, Xaq Pitkow, Dora E. Angelaki
Real world choices often involve balancing decisions that are optimized for the short-vs. long-term. Here, we reason that apparently sub-optimal single trial decisions in macaques may in fact reflect long-term, strategic planning. We demonstrate that macaques freely navigating in VR for sequentially presented targets…
Ruiyi Zhang, Xaq Pitkow, Dora E. Angelaki
The brain may have evolved a modular architecture for daily tasks, with circuits featuring functionally specialized modules that match the task structure. We hypothesize that this architecture enables better learning and generalization than architectures with less specialized modules. To test this, we trained…
Noeline W. Prins, Justin C. Sanchez, Abhishek Prasad
Brain-Machine Interfaces (BMIs) can be used to restore function in people living with paralysis. Current BMIs require extensive calibration that increase the set-up times and external inputs for decoder training that may be difficult to produce in paralyzed individuals. Both these factors have presented challenges in…
Elena Zamaraeva, Christopher M. Collins, Dmytro Antypov, Vladimir V. Gusev + 6 more
Crystal Structure Prediction (CSP) is a fundamental computational problem in materials science. Basin-hopping is a prominent CSP method that combines global Monte Carlo sampling to search over candidate trial structures with local energy minimisation of these candidates. The sampling uses a stochastic policy to…
Authors not listed
Computer-aided synthesis planning aims to identify viable synthetic routes from a target compound to readily available building blocks by iteratively decomposing molecules into smaller precursors. Self-play search algorithms, trained with simulated experience, reach state-of-the-art performance. However, these methods…
Etinosa Osaro, Yamil Colón
The application of machine learning (ML) techniques in materials science has revolutionized the pace and scope of materials research and design. In the case of metal-organic frameworks (MOFs), a promising class of materials due to their tunable properties and versatile applications in gas adsorption and separation, ML…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…