20 papers · ranked by Valyu relevance
Yoshimasa Kubo, Eric Chalmers, Artur Luczak
Backpropagation has been used to train neural networks for many years, allowing them to solve a wide variety of tasks like image classification, speech recognition, and reinforcement learning tasks. But the biological plausibility of backpropagation as a mechanism of neural learning has been questioned. Equilibrium…
Yoshimasa Kubo, Eric Chalmers, Artur Luczak
Backpropagation (BP) has been used to train neural networks for many years, allowing them to solve a wide variety of tasks like image classification, speech recognition, and reinforcement learning tasks. But the biological plausibility of BP as a mechanism of neural learning has been questioned. Equilibrium Propagation…
Liyuan Zheng, Tanner Fiez, Zane Alumbaugh, Benjamin J. Chasnov + 1 more
'Lillian J. Ratliff'] The hierarchical interaction between the actor and critic in actor-critic based reinforcement learning algorithms naturally lends itself to a game-theoretic interpretation. We adopt this viewpoint and model the actor and critic interaction as a two-player general-sum game with a leader-follower…
Shalabh Bhatnagar, Vivek S. Borkar, Soumyajit Guin
We revisit the standard formulation of tabular actor-critic algorithm as a two time-scale stochastic approximation with value function computed on a faster time-scale and policy computed on a slower time-scale. This emulates policy iteration. We observe that reversal of the time scales will in fact emulate value…
Simone Parisi, Voot Tangkaratt, Jan Peters, Mohammad Emtiyaz Khan
Actor-critic methods can achieve incredible performance on difficult reinforcement learning problems, but they are also prone to instability. This is partly due to the interaction between the actor and the critic during learning, e.g., an inaccurate step taken by one of them might adversely affect the other and…
Charline Tessereau, Reuben O’Dea, Stephen Coombes, Tobias Bast
Humans and non-human animals show great flexibility in spatial navigation, including the ability to return to specific locations based on as few as one single experience. To study spatial navigation in the laboratory, watermaze tasks, in which rats have to find a hidden platform in a pool of cloudy water surrounded by…
Junfeng Wen, Saurabh Kumar, Ramki Gummadi, Dale Schuurmans
Actor-critic (AC) methods are ubiquitous in reinforcement learning. Although it is understood that AC methods are closely related to policy gradient (PG), their precise connection has not been fully characterized previously. In this paper, we explain the gap between AC and PG methods by identifying the exact adjustment…
Sukriti Verma, Ayush Chopra, Jayakumar Subramanian, Mausoom Sarkar + 3 more
'Nikaash Puri' 'Piyush Gupta' 'Balaji Krishnamurthy'] The two-time scale nature of SAC, which is an actor-critic algorithm, is characterised by the fact that the critic estimate has not converged for the actor at any given time, but since the critic learns faster than the actor, it ensures eventual consistency between…
Anton Wiehe, Nil Stolt-Ansó, Mădălina M. Drugan, Marco Wiering
In this paper, a new offline actor-critic learning algorithm is introduced: Sampled Policy Gradient (SPG). SPG samples in the action space to calculate an approximated policy gradient by using the critic to evaluate the samples. This sampling allows SPG to search the action-Q-value space more globally than…
Menghao Wu, Yanbin Gao, Alexander Jung, Qiang Zhang + 1 more
Model-free reinforcement learning is a powerful and efficient machine-learning paradigm which has been generally used in the robotic control domain. In the reinforcement learning setting, the value function method learns policies by maximizing the state-action value (Q value), but it suffers from inaccurate Q…
Sharan Vaswani, Amirreza Kazemi, Reza Babanezhad, Nicolas Le Roux
Actor-critic (AC) methods are widely used in reinforcement learning (RL), and benefit from the flexibility of using any policy gradient method as the actor and value-based method as the critic. The critic is usually trained by minimizing the TD error, an objective that is potentially decorrelated with the true goal of…
Anushka Deshpande
The aim of this paper is twofold. First, it seeks to uncover the algorithms that humans and other animals employ for learning in decision-making strategies within non-zero-sum games, specifically focusing on fully observable iterated prisoner’s dilemma scenarios. Second, it aims to develop a new model to explain…
Massimo Silvetti, Eliana Vassena, Elger Abrahamse, Tom Verguts
The dorsal anterior cingulate cortex (dACC) is central in higher-order cognition and behavioural flexibility. The computational nature of this region, however, has remained elusive. Here we propose a new model – the Reinforcement Meta Learner (RML) – based on the bidirectional anatomical connections of the ACC with…
Haifei Zhang, Jian Xu, Jian Zhang, Quan Liu
The traditional Deep Deterministic Policy Gradient (DDPG) algorithm has been widely used in continuous action spaces, but it still suffers from the problems of easily falling into local optima and large error fluctuations. Aiming at these deficiencies, this paper proposes a dual-actor-dual-critic DDPG algorithm…
Babak Mahmoudi, Justin C. Sanchez, Josh Bongard
Background In the development of Brain Machine Interfaces (BMIs), there is a great need to enable users to interact with changing environments during the activities of daily life. It is expected that the number and scope of the learning tasks encountered during interaction with the environment as well as the pattern of…
Cesar Guevara, Bilal Alatas
Currently, the stock market is attractive, and it is challenging to develop an efficient investment model with high accuracy due to changes in the values of the shares for political, economic, and social reasons. This article presents an innovative proposal for a short-term, automatic investment model to reduce capital…
Ruiyi Zhang, Xaq Pitkow, Dora E. Angelaki
The brain may have evolved a modular architecture for daily tasks, with circuits featuring functionally specialized modules that match the task structure. We hypothesize that this architecture enables better learning and generalization than architectures with less specialized modules. To test this, we trained…
Jean-Paul Noel, Ruiyi Zhang, Xaq Pitkow, Dora E. Angelaki
Real world choices often involve balancing decisions that are optimized for the short-vs. long-term. Here, we reason that apparently sub-optimal single trial decisions in macaques may in fact reflect long-term, strategic planning. We demonstrate that macaques freely navigating in VR for sequentially presented targets…
Arash Khodadadi, Pegah Fakhari, Jerome R. Busemeyer
When animals have to make a number of decisions during a limited time interval, they face a fundamental problem: how much time they should spend on each decision in order to achieve the maximum possible total outcome. Deliberating more on one decision usually leads to more outcome but less time will remain for other…
Pranav Mahajan, Ben Seymour
The seminal reward prediction error theory of dopamine function faces several key challenges. Most notable is the difficulty learning multiple rewards simultaneously, inefficient on-policy learning, and accounting for heterogeneous striatal responses in the tail of the striatum. We propose a normative framework, based…