Search · four archives
Search · four archives
24 papers · ranked by Valyu relevance
Malcolm G. Campbell, Yongsoo Ra, Zhiqin Chen, Shudi Xu + 4 more
The neurotransmitter dopamine plays a major role in learning by acting as a teaching signal to update the brain’s predictions about rewards. A leading theory proposes that this process is analogous to a reinforcement learning algorithm called temporal difference (TD) learning, and that dopamine acts as the error term…
Wiebke Potjans, Markus Diesmann, Abigail Morrison, Tim Behrens
An open problem in the field of computational neuroscience is how to link synaptic plasticity to system-level learning. A promising framework in this context is temporal-difference (TD) learning. Experimental evidence that supports the hypothesis that the mammalian brain performs temporal-difference learning includes…
Mokhaled N. A. Al-Hamadani, Mohammed A. Fadhel, Laith Alzubaidi, Harangi Balazs + 2 more
'Harangi Balazs' 'Antonio Fernández-Caballero' 'Dominique Gruyer'] Reinforcement learning (RL) has emerged as a dynamic and transformative paradigm in artificial intelligence, offering the promise of intelligent decision-making in complex and dynamic environments. This unique feature enables RL to address sequential…
Kristopher T. Jensen
Reinforcement learning has a rich history in neuroscience, from early work on dopamine as a reward prediction error signal for temporal difference learning (Schultz et al., 1997) to recent work suggesting that dopamine could implement a form of 'distributional reinforcement learning' popularized in deep learning…
Zeb Kurth-Nelson, A. David Redish, Olaf Sporns
Temporal-difference (TD) algorithms have been proposed as models of reinforcement learning (RL). We examine two issues of distributed representation in these TD algorithms: distributed representations of belief and distributed discounting factors. Distributed representation of belief allows the believed state of the…
Beren Millidge, Mark Walton, Rafal Bogacz
An influential theory posits that dopaminergic neurons in the mid-brain implement a model-free reinforcement learning algorithm based on temporal difference (TD) learning. A fundamental assumption of this model is that the reward function being optimized is fixed. However, for biological creatures the ‘reward function’…
Richard S. Sutton, B. K. Tanner
We introduce a generalization of temporal-difference (TD) learning to networks of interrelated predictions. Rather than relating a single prediction to itself at a later time, as in conventional TD methods, a TD network relates each prediction in a set of predictions to other predictions in the set at a later time. TD…
Farnaz Adib Yaghmaie, Lennart Ljung
The emerging field of Reinforcement Learning (RL) has led to impressive results in varied domains like strategy games, robotics, etc. This handout aims to give a simple introduction to RL from control perspective and discuss three possible approaches to solve an RL problem: Policy Gradient, Policy Iteration, and…
Benton Girdler, William Caldbeck, Jihye Bae
Creating flexible and robust brain machine interfaces (BMIs) is currently a popular topic of research that has been explored for decades in medicine, engineering, commercial, and machine-learning communities. In particular, the use of techniques using reinforcement learning (RL) has demonstrated impressive results but…
Vitchyr H. Pong, Shixiang Gu, Murtaza Dalal, Sergey Levine
Model-free reinforcement learning (RL) is a powerful, general tool for learning complex behaviors. However, its sample efficiency is often impractically large for solving challenging real-world problems, even with off-policy algorithms such as Q-learning. A limiting factor in classic model-free RL is that the learning…
Venkata S Aditya Tarigoppula, John S Choi, John P Hessburg, David B McNiel + 2 more
Temporal difference reinforcement learning (TDRL) accurately models associative learning observed in animals, where they learn to associate outcome predicting environmental states, termed conditioned stimuli (CS), with the value of outcomes, such as rewards, termed unconditioned stimuli (US). A component of TDRL is the…
Can Demircan, Tankred Saanum, Akshay K. Jagadish, Marcel Binz + 1 more
Language Models Authors: ['Can Demircan' 'Tankred Saanum' 'Akshay K. Jagadish' 'Marcel Binz' 'Eric Schulz'] In-context learning, the ability to adapt based on a few examples in the input prompt, is a ubiquitous feature of large language models (LLMs). However, as LLMs' in-context learning abilities continue to improve…
Kim T. Blackwell, Kenji Doya, Ming Bo Cai
A major advance in understanding learning behavior stems from experiments showing that reward learning requires dopamine inputs to striatal neurons and arises from synaptic plasticity of cortico-striatal synapses. Numerous reinforcement learning models mimic this dopamine-dependent synaptic plasticity by using the…
Ian Cone, Claudia Clopath, Harel Z. Shouval
The dominant theoretical framework to account for reinforcement learning in the brain is temporal difference (TD) reinforcement learning. The TD framework predicts that some neuronal elements should represent the reward prediction error (RPE), which means they signal the difference between the expected future rewards…
Shivam Kalhan, Marta I. Garrido, Robert Hester, A. David Redish
Dysfunction in learning and motivational systems are thought to contribute to addictive behaviours. Previous models have suggested that dopaminergic roles in learning and motivation could produce addictive behaviours through pharmacological manipulations that provide excess dopaminergic signalling towards these…
Eric Chalmers, Santina Duarte, Xena Al-Hejji, Daniel Devoe + 2 more
Deep Reinforcement Learning is a branch of artificial intelligence that uses artificial neural networks to model reward-based learning as it occurs in biological agents. Here we modify a Deep Reinforcement Learning approach by imposing a suppressive effect on the connections between neurons in the artificial network -…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Vincent Graziano, F. Gómez, Mark Ring, Juergen Schmidhuber
Traditional Reinforcement Learning (RL) has focused on problems involving many states and few actions, such as simple grid worlds. Most real world problems, however, are of the opposite type, Involving Few relevant states and many actions. For example, to return home from a conference, humans identify only few subgoal…
Pranav Mahajan, Ben Seymour
The seminal reward prediction error account of dopamine has been highly successful, but faces several key challenges. Most notable are the difficulty of learning multiple rewards simultaneously, inefficient on-policy learning, and accounting for the heterogeneous striatal responses observed across and within striatal…
Victor Geadah, Jonathan W. Pillow
Identifying the learning rules that govern behavior is a central problem in neuroscience. While reinforcement learning (RL) offers a unifying theoretical framework, most empirical studies of animal learning behavior have focused on non-stationary environments (e.g. changing reward probabilities in a known task), as…
Authors not listed
Accurate prediction of chemical reaction yields remains essential for accelerating synthesis optimization, yet current machine learning models face critical limitations in capturing temporal dynamics, providing calibrated uncertainty estimates, and explicitly modeling reactant-to-product transformations. Here we…
Etinosa Osaro, Yamil Colón
The application of machine learning (ML) techniques in materials science has revolutionized the pace and scope of materials research and design. In the case of metal-organic frameworks (MOFs), a promising class of materials due to their tunable properties and versatile applications in gas adsorption and separation, ML…
Authors not listed
Accurately modeling the dynamics of open quantum systems is critical for advancing quantum technologies, yet traditional methods often struggle with balancing accuracy and efficiency. Machine learning (ML) offers a promising alternative, particularly through recursive models that predict system evolution based on the…
Jeff Guo, Vendy Fialková, Juan Diego Arango, Christian Margreitter + 4 more
Reinforcement learning (RL) is a powerful paradigm that has gained popularity across multiple domains. However, applying RL may come at a cost of multiple interactions between the agent and the environment. This cost can be especially pronounced when the single feedback from the environment is slow or computationally…