21 papers · ranked by Valyu relevance
Tom Schaul, John Quan, Ioannis Antonoglou, David Silver
Experience replay lets online reinforcement learning agents remember and reuse experiences from the past. In prior work, experience transitions were uniformly sampled from a replay memory. However, this approach simply replays transitions at the same frequency that they were originally experienced, regardless of their…
Thommen George Karimpanal, Roland Bouffanais
Experience replay is one of the most commonly used approaches to improve the sample efficiency of reinforcement learning algorithms. In this work, we propose an approach to select and replay sequences of transitions in order to accelerate the learning of a reinforcement learning agent in an off-policy setting. In…
Tyler L. Hayes, Giri P. Krishnan, Maxim Bazhenov, Hava T. Siegelmann + 2 more
'Terrence J. Sejnowski' 'Christopher Kanan'] Replay is the reactivation of one or more neural patterns, which are similar to the activation patterns experienced during past waking experiences. Replay was first observed in biological neural networks during sleep, and it is now thought to play a critical role in memory…
Daochen Zha, Kwei-Herng Lai, Kaixiong Zhou, Xia Hu
Experience replay enables reinforcement learning agents to memorize and reuse past experiences, just as humans replay memories for the situation at hand. Contemporary off-policy algorithms either replay past experiences uniformly or utilize a rulebased replay strategy, which may be sub-optimal. In this work, we…
Zhenglong Zhou, Michael J. Kahana, Anna C. Schapiro
During rest and sleep, sequential neural activation patterns corresponding to awake experience re-emerge, and this replay has been shown to benefit subsequent behavior and memory. Whereas some studies show that replay directly recapitulates recent experience, others demonstrate that replay systematically deviates from…
Zhenglong Zhou, Michael J. Kahana, Anna C. Schapiro
Replay in the brain is not a simple recapitulation of recent experience, with awake replay often unrolling in reverse temporal order upon receipt of reward, in a manner dependent on reward magnitude. These findings have led to the proposal that replay is optimized for learning value-based predictions in accordance with…
Elisa Massi, Jeanne Barthélemy, Juliane Mailly, Rémi Dromnelle + 4 more
'Julien Canitrot' 'Esther Poniatowski' 'Benoît Girard' 'Mehdi Khamassi'] Experience replay is widely used in AI to bootstrap reinforcement learning (RL) by enabling an agent to remember and reuse past experiences. Classical techniques include shuffled-, reversed-ordered- and prioritized-memory buffers, which have…
Zhenglong Zhou, Michael J Kahana, Anna C Schapiro, Mimi Liljeholm + 1 more
During rest and sleep, sequential neural activation patterns corresponding to awake experience re-emerge, and this replay has been shown to benefit subsequent behavior and memory. Whereas some studies show that replay directly recapitulates recent experience, others demonstrate that replay systematically deviates from…
Emma L. Roscow, Raymond Chua, Rui Ponte Costa, Matthew W. Jones + 1 more
'Nathan F. Lepora'] 1Centre de Recerca Matemàtica, Bellaterra, Spain; 2McGill University and Mila, Montréal, Canada; 3Bristol Computational Neuroscience Unit, Intelligent Systems Lab, Department of Computer Science, University of Bristol, UK; 4School of Physiology, Pharmacology and Neuroscience, University of Bristol…
Yunhan Lin, Zhijie Zhang, Yijian Tan, Hao Fu + 1 more
To address the challenges of sample utilization efficiency and managing temporal dependencies, this paper proposes an efficient path planning method for mobile robot in dynamic environments based on an improved twin delayed deep deterministic policy gradient (TD3) algorithm. The proposed method, named PL-TD3…
Jeff Guo, Philippe Schwaller
Sample efficiency is a fundamental challenge in de novo molecular design. Ideally, molecular generative models should learn to satisfy desired objectives under minimal oracle evaluations (computational prediction or wet-lab experiment). This problem becomes more apparent when using oracles that can provide increased…
Jeff Guo, Philippe Schwaller
Sample efficiency is a fundamental challenge in de novo molecular design. Ideally, molecular generative models should learn to satisfy desired objectives under minimal oracle evaluations (computational prediction or wet-lab experiment). This problem becomes more apparent when using oracles that can provide increased…
Yizhi Yuan, Marcelo G. Mattar
Prioritized experience replay is a reinforcement learning technique whereby agents speed up learning by replaying useful past experiences. This usefulness is quantified as the expected gain from replaying the experience, a quantity often approximated as the prediction error (TD-error). However, recent work in…
Liran Szlak, Ohad Shamir
Experience replay [Lin, 1993, Mnih et al., 2015] is a widely used technique to achieve efficient use of data and improved performance in RL algorithms. In experience replay, past transitions are stored in a memory buffer and re-used during learning. Various suggestions for sampling schemes from the replay buffer have…
Egor Rotinov
This paper describes an improvement in Deep Q-learning called Reverse Experience Replay (also RER) that solves the problem of sparse rewards and helps to deal with reward maximizing tasks by sampling transitions successively in reverse order. On tasks with enough experience for training and enough Experience Replay…
Daniel N. Barry, Bradley C. Love
Replay can consolidate memories through offline neural reactivation related to past experiences. Category knowledge is learned across multiple experiences, and its subsequent generalisation is promoted by consolidation and replay during rest and sleep. However, aspects of replay are difficult to determine from…
Authors not listed
Large Language Models (LLMs) based on transformer architectures excel at internet-scale tasks. However, real-world scientific scenarios—such as synthetic chemistry laboratories and autonomous experimental setups—typically involve incremental data generation in batches as new chemical reactions are conducted, unlike…
Yotam Sagiv, Thomas Akam, Ilana B. Witten, Nathaniel D. Daw
Although hippocampal place cells replay nonlocal trajectories, the computational function of these events remains controversial. One hypothesis, formalized in a prominent reinforcement learning account, holds that replay plans routes to current goals. However, recent puzzling data appear to contradict this perspective…
Daniel N Barry, Bradley C Love
Replay can consolidate memories through offline neural reactivation related to past experiences. Category knowledge is learned across multiple experiences, and its subsequent generalization is promoted by consolidation and replay during rest and sleep. However, aspects of replay are difficult to determine from…
Elliot A. Ludvig, Mahdieh S. Mirian, E. James Kehoe, Richard S. Sutton
We develop an extension of the Rescorla-Wagner model of associative learning. In addition to learning from the current trial, the new model supposes that animals store and replay previous trials, learning from the replayed trials using the same learning rule. This simple idea provides a unified explanation for diverse…
Georgy Antonov, Christopher Gagne, Eran Eldar, Peter Dayan
The replay of task-relevant trajectories is known to contribute to memory consolidation and improved task performance. A wide variety of experimental data show that the content of replayed sequences is highly specific and can be modulated by reward as well as other prominent task variables. However, the rules governing…