24 papers · ranked by Valyu relevance
Ghoshana Bista
Deep reinforcement learning has evolved from classical dynamic programming, temporal-difference learning, and tabular control into a broad framework for sequential decision-making under uncertainty. This book provides a structured introduction to that evolution, emphasizing not only how reinforcement learning…
Perkins, Daniel, Escobar, Oscar J. + 2 more
We present a detailed study of Deep Q-Networks in finite environments, emphasizing the impact of epsilon-greedy exploration schedules and prioritized experience replay. Through systematic experimentation, we evaluate how variations in epsilon decay schedules affect learning efficiency, convergence behavior, and reward…
Anil Kumar Yadav, Purushottam Sharma, Xiaochun Cheng, Shiv Shankar Prasad Shukla
Path selection and planning are crucial for autonomous mobile robots (AMRs) to navigate efficiently and avoid obstacles. Traditional methods rely on analytical search to identify the shortest distance. However, Reinforcement learning enhances performance by optimizing a sequence of actions efficiently. It is an…
Zhizuo Chen, Theodore T. Allen
Algorithms developed under stationary Markov Decision Processes (MDPs) often face challenges in non-stationary environments, and infinite-horizon formulations may not directly apply to finite-horizon tasks. To address these limitations, we introduce the Non-stationary and Varying-discounting MDP (NVMDP) framework…
Abhijit Sen, Sonali Panda, Mahima Arya, Subhajit Patra + 2 more
Abhijit Sen , 1, ∗ Sonali Panda , 2, † Mahima Arya , 1, ‡ Subhajit Patra , 3, § Zizhan Zheng , 4, ¶ and Denys I. Bondar 1, ∗∗ 1 Department of Physics and Engineering Physics, Tulane University, New Orleans, LA 70118, USA 2 Department of Physics, Indian Institute of Technology, Dhanbad, India 3 Department of Electrical…
Wang, Likun, Zhang, Xiangteng + 12 more
Exploration is fundamental to reinforcement learning (RL), as it determines how effectively an agent discovers and exploits the underlying structure of its environment to achieve optimal performance. Existing exploration methods generally fall into two categories: active exploration and passive exploration. The former…
Pranav Mahajan, Ben Seymour
The seminal reward prediction error theory of dopamine function faces several key challenges. Most notable is the difficulty learning multiple rewards simultaneously, inefficient on-policy learning, and accounting for heterogeneous striatal responses in the tail of the striatum. We propose a normative framework, based…
Pranav Mahajan, Ben Seymour
The seminal reward prediction error account of dopamine has been highly successful, but faces several key challenges. Most notable are the difficulty of learning multiple rewards simultaneously, inefficient on-policy learning, and accounting for the heterogeneous striatal responses observed across and within striatal…
Armin Bazarjani, Payam Piray
Cognitive maps enable flexible behavior by providing reusable internal representations of task structure. The successor representation, a predictive map that encodes expected future state occupancy, has been proposed as one way such maps might be computed in the brain, but its policy dependence severely limits flexible…
Nicholas Zolman, Christian Lagemann, Urban Fasel, J. Nathan Kutz + 1 more
Deep reinforcement learning (DRL) has shown significant promise for uncovering sophisticated control policies that interact in complex environments, such as stabilizing a tokamak fusion reactor or minimizing the drag force on an object in a fluid flow. However, DRL requires an abundance of training examples and may…
Markus D. Solbach, John K. Tsotsos
Reinforcement Learning is a mature technology, often suggested as a potential route towards Artificial General Intelligence, with the ambitious goal of replicating the wide range of abilities found in natural and artificial intelligence, including the complexities of human cognition. While RL had shown successes in…
Francesca Greenstreet, Jesse P. Geerts, Juan A. Gallego, Claudia Clopath
The initial stage of learning motor skills involves exploring vast action spaces, making it impractical to learn the value of every possible action independently. This poses a challenge for standard reinforcement learning approaches, which excel in constrained domains but struggle when the space of possible actions is…
Authors not listed
Realizing the promise of artificial intelligence (AI) to accelerate scientific progress and deliver technological impact depends on how effectively AI can be integrated into real-world decision- making processes. As Peter Norvig states, “Somewhat remarkably, almost all AI research until very recently has assumed that…
Amar Ahmad, Yvonne Vallès, Youssef Idaghdour
Reinforcement learning (RL) has achieved remarkable success in controlled environments, demonstrating superhuman performance in domains such as game playing and simulated robotics. However, its transition to real-world applications remains constrained by fundamental statistical challenges that limit scalability…
Emma L. Roscow, Timothy Howe, Nathan F. Lepora, Matthew W. Jones
Neural activity encoding recent experiences is replayed during sleep and rest to promote consolidation of memories. However, precisely which features of experience influence replay prioritisation to optimise adaptive behaviour remains unclear. Here, we trained adult male rats on a novel maze-based reinforcement…
Authors not listed
The integration of machine learning methods is transforming many areas of research by, for instance, accelerating molecular dynamics simulations and enabling improved prediction and optimization of chemical reactions. However, despite this progress, the adoption of data-driven approaches in atomic layer deposition…
Nicolas Diekmann, Silke Lissek, Metin Üngör, Sen Cheng
The progress of learning is usually quantified by averaging responses across participants and/or multiple trials within a block. However, such approaches obscure the trial-by-trial progress of learning, which has been shown recently to express a rich variety of dynamics. An alternative approach which does not suffer…
Wei Yin, Sanad H. Ragab, Michael G. Tyshenko, Teresa Feria Arroyo + 2 more
Several machine learning (ML) and deep learning (DL) methods have been used to predict the presence of species in classification problems. Another set of methods, called reinforcement learning (RL), has been used in training agents to perform various tasks, but not in predicting species distribution. Culex pipiens…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Christina E. Wierenga, Carina S. Brown, Erin E. Reilly
Purpose of Review We review recent literature on instrumental reinforcement learning involving decision-making in anorexia nervosa (AN) to understand mechanisms underlying symptoms of AN, such as rigid pursuit of weight loss despite negative consequences. Recent Findings Relatively consistent findings indicate worse…
Francesco Rigoli
Influential cognitive science theories postulate that decision-making is based on treating expected outcomes as incentives according to a reward function. Yet a systematic analysis of the learning processes that determine the reward function remains to be carried out. The paper fills this gap by examining the…
Gaspard Lambrechts, Adrien Bolland, Daniel Ebi, Damien Ernst
Much like humans benefit from guidance while learning, reinforcement learning algorithms may benefit from additional supervision beyond rewards. Leveraging additional information during training to learn better representations and behaviors has been the focus of asymmetric reinforcement learning. This learning paradigm…
Authors not listed
Inverse molecular design aims to generate novel chemical structures that satisfy multiple property constraints, yet reinforcement-learning (RL) fine-tuning can be sensitive to how objectives are converted into a scalar reward. Here, we systematically analyze how scalarization choices and stabilization mechanisms shape…
Authors not listed
Three-dimensional molecular generative models have emerged that produce de novo molecules both unconditionally and conditionally, e.g., within protein pockets. However, steering those models in a specific region of the chemical space that satisfies a set of desired properties remains challenging. In this study, we…