20 papers · ranked by Valyu relevance
Malcolm G. Campbell, Yongsoo Ra, Zhiqin Chen, Shudi Xu + 4 more
The neurotransmitter dopamine plays a major role in learning by acting as a teaching signal to update the brain’s predictions about rewards. A leading theory proposes that this process is analogous to a reinforcement learning algorithm called temporal difference (TD) learning, and that dopamine acts as the error term…
Ian Cone, Claudia Clopath, Harel Z. Shouval
The dominant theoretical framework to account for reinforcement learning in the brain is temporal difference learning (TD) learning, whereby certain units signal reward prediction errors (RPE). The TD algorithm has been traditionally mapped onto the dopaminergic system, as firing properties of dopamine neurons can…
Han-Dong Lim, Donghwan Lee
Off-policy learning ability is an important feature of reinforcement learning (RL) for practical applications. However, even one of the most elementary RL algorithms, temporal-difference (TD) learning, is known to suffer form divergence issue when the off-policy scheme is used together with linear function…
Mokhaled N. A. Al-Hamadani, Mohammed A. Fadhel, Laith Alzubaidi, Harangi Balazs + 2 more
'Harangi Balazs' 'Antonio Fernández-Caballero' 'Dominique Gruyer'] Reinforcement learning (RL) has emerged as a dynamic and transformative paradigm in artificial intelligence, offering the promise of intelligent decision-making in complex and dynamic environments. This unique feature enables RL to address sequential…
Ian Cone, Claudia Clopath, Harel Z. Shouval
Dopamine (DA) releasing neurons in the midbrain learn response patterns that represent reward prediction error (RPE). Typically, models proposing a mechanistic explanation for how dopamine neurons learn to exhibit RPE are based on temporal difference (TD) learning, a machine learning algorithm. However, mechanistic…
Rong J. B. Zhu, James M. Murray
Off-policy algorithms, in which a behavior policy differs from the target policy and is used to gain experience for learning, have proven to be of great practical value in reinforcement learning. However, even for simple convex problems such as linear value function approximation, these algorithms are not guaranteed to…
Kristopher T. Jensen
Reinforcement learning has a rich history in neuroscience, from early work on dopamine as a reward prediction error signal for temporal difference learning (Schultz et al., 1997) to recent work suggesting that dopamine could implement a form of 'distributional reinforcement learning' popularized in deep learning…
Anthony M.V. Jakob, John G. Mikhael, Allison E. Hamilos, John A. Assad + 1 more
The role of dopamine as a reward prediction error signal in reinforcement learning tasks has been well-established over the past decades. Recent work has shown that the reward prediction error interpretation can also account for the effects of dopamine on interval timing by controlling the speed of subjective time.…
Kim T. Blackwell, Kenji Doya, Ming Bo Cai
A major advance in understanding learning behavior stems from experiments showing that reward learning requires dopamine inputs to striatal neurons and arises from synaptic plasticity of cortico-striatal synapses. Numerous reinforcement learning models mimic this dopamine-dependent synaptic plasticity by using the…
Brett Daley, Marlos C. Machado, Martha White
The recency heuristic in reinforcement learning is the assumption that stimuli that occurred closer in time to an acquired reward should be more heavily reinforced. The recency heuristic is one of the key assumptions made by TD(λ), which reinforces recent experiences according to an exponentially decaying weighting. In…
Eric Chalmers, Artur Luczak
Developments in reinforcement learning (RL) have allowed algorithms to achieve impressive performance in highly complex, but largely static problems. In contrast, biological learning seems to value efficiency of adaptation to a constantly-changing world. Here we build on a recently-proposed neuronal learning rule that…
Brittany Liebenow, Rachel Jones, Emily DiMarco, Jonathan D. Trattner + 7 more
'Joseph Humphries' 'L. Paul Sands' 'Kasey P. Spry' 'Christina K. Johnson' 'Evelyn B. Farkas' 'Angela Jiang' 'Kenneth T. Kishida'] In the DSM-5, psychiatric diagnoses are made based on self-reported symptoms and clinician-identified signs. Though helpful in choosing potential interventions based on the available…
Paul Masset, Pablo Tano, HyungGoo R. Kim, Athar N. Malik + 2 more
To thrive in complex environments, animals and artificial agents must learn to act adaptively to maximize fitness and rewards. Such adaptive behavior can be learned through reinforcement learning^1^, a class of algorithms that has been successful at training artificial agents^2–6^ and at characterizing the firing of…
Xiang Gao, Tianyuan Liu, Yisha Li, Jingxin Liu + 4 more
With the rapid advancement of Transformer-based Large Language Models (LLMs), generative recommendation has shown great potential in enhancing both the accuracy and semantic understanding of modern recommender systems. Compared to LLMs, the Decision Transformer (DT) is a lightweight generative model applied to…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Emily K. DiMarco, Ashley Ratcliffe Shipp, Kenneth T. Kishida
Time perception is often investigated in animal models and in humans using instrumental paradigms where reinforcement learning (RL) and associated dopaminergic processes have modulatory effects. For example, interval timing, which includes the judgment of relatively short intervals of time (ranging from milliseconds to…
Patrick M. Pilarski, Andrew Butcher, Elnaz Davoodi, Michael Bradley Johanson + 6 more
'Michael Bradley Johanson' 'Dylan J. A. Brenneis' 'Adam S. R. Parker' 'Leslie Acker' 'Matthew Botvinick' 'Joseph Modayil' 'Adam White'] Patrick M. Pilarski1,3,4, Andrew Butcher1 , Elnaz Davoodi1 , Michael Bradley Johanson1 , Dylan J. A. Brenneis1 , Adam S. R. Parker3,4, Leslie Acker1 , Matthew M. Botvinick2 , Joseph…
Margarida Sousa, Pawel Bujalski, Bruno F. Cruz, Kenway Louie + 2 more
Learning to predict rewards is a fundamental driver of adaptive behavior. Midbrain dopamine neurons (DANs) play a key role in such learning by signaling reward prediction errors (RPEs) that teach recipient circuits about expected rewards given current circumstances and actions. However, the algorithm that DANs are…
Authors not listed
Accurate prediction of chemical reaction yields remains essential for accelerating synthesis optimization, yet current machine learning models face critical limitations in capturing temporal dynamics, providing calibrated uncertainty estimates, and explicitly modeling reactant-to-product transformations. Here we…
David Mathar, Annika Wiebe, Deniz Tuzsus, Kilian Knauth + 1 more
Computational psychiatry focuses on identifying core cognitive processes that appear altered across a broad range of psychiatric disorders. Temporal discounting of future rewards and model-based control during reinforcement learning have proven as two promising candidates. Despite its trait-like stability, temporal…