22 papers · ranked by Valyu relevance
Esther Mondragón, Jonathan Gray, Eduardo Alonso, Charlotte Bonardi + 2 more
This paper presents a novel representational framework for the Temporal Difference (TD) model of learning, which allows the computation of configural stimuli - cumulative compounds of stimuli that generate perceptual emergents known as configural cues. This Simultaneous and Serial Configural-cue Compound Stimuli…
Vitchyr H. Pong, Shixiang Gu, Murtaza Dalal, Sergey Levine
Model-free reinforcement learning (RL) is a powerful, general tool for learning complex behaviors. However, its sample efficiency is often impractically large for solving challenging real-world problems, even with off-policy algorithms such as Q-learning. A limiting factor in classic model-free RL is that the learning…
Malcolm G. Campbell, Yongsoo Ra, Zhiqin Chen, Shudi Xu + 4 more
The neurotransmitter dopamine plays a major role in learning by acting as a teaching signal to update the brain’s predictions about rewards. A leading theory proposes that this process is analogous to a reinforcement learning algorithm called temporal difference (TD) learning, and that dopamine acts as the error term…
Ian Cone, Claudia Clopath, Harel Z. Shouval
The dominant theoretical framework to account for reinforcement learning in the brain is temporal difference learning (TD) learning, whereby certain units signal reward prediction errors (RPE). The TD algorithm has been traditionally mapped onto the dopaminergic system, as firing properties of dopamine neurons can…
Wiebke Potjans, Markus Diesmann, Abigail Morrison, Tim Behrens
An open problem in the field of computational neuroscience is how to link synaptic plasticity to system-level learning. A promising framework in this context is temporal-difference (TD) learning. Experimental evidence that supports the hypothesis that the mammalian brain performs temporal-difference learning includes…
Zeb Kurth-Nelson, A. David Redish, Olaf Sporns
Temporal-difference (TD) algorithms have been proposed as models of reinforcement learning (RL). We examine two issues of distributed representation in these TD algorithms: distributed representations of belief and distributed discounting factors. Distributed representation of belief allows the believed state of the…
Richard S. Sutton, B. K. Tanner
We introduce a generalization of temporal-difference (TD) learning to networks of interrelated predictions. Rather than relating a single prediction to itself at a later time, as in conventional TD methods, a TD network relates each prediction in a set of predictions to other predictions in the set at a later time. TD…
Kristopher T. Jensen
Reinforcement learning has a rich history in neuroscience, from early work on dopamine as a reward prediction error signal for temporal difference learning (Schultz et al., 1997) to recent work suggesting that dopamine could implement a form of 'distributional reinforcement learning' popularized in deep learning…
Kim T. Blackwell, Kenji Doya, Ming Bo Cai
A major advance in understanding learning behavior stems from experiments showing that reward learning requires dopamine inputs to striatal neurons and arises from synaptic plasticity of cortico-striatal synapses. Numerous reinforcement learning models mimic this dopamine-dependent synaptic plasticity by using the…
Paul Masset, Pablo Tano, HyungGoo R. Kim, Athar N. Malik + 2 more
To thrive in complex environments, animals and artificial agents must learn to act adaptively to maximize fitness and rewards. Such adaptive behavior can be learned through reinforcement learning^1^, a class of algorithms that has been successful at training artificial agents^2–6^ and at characterizing the firing of…
Markus Dumke
Temporal-difference (TD) learning is an important field in reinforcement learning. Sarsa and Q-Learning are among the most used TD algorithms. The Q(σ) algorithm (Sutton and Barto (2017)) unifies both. This paper extends the Q(σ) algorithm to an online multi-step algorithm Q(σ, λ) using eligibility traces and…
Xiang Gao, Tianyuan Liu, Yisha Li, Jingxin Liu + 4 more
With the rapid advancement of Transformer-based Large Language Models (LLMs), generative recommendation has shown great potential in enhancing both the accuracy and semantic understanding of modern recommender systems. Compared to LLMs, the Decision Transformer (DT) is a lightweight generative model applied to…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Emily K. DiMarco, Ashley Ratcliffe Shipp, Kenneth T. Kishida
Time perception is often investigated in animal models and in humans using instrumental paradigms where reinforcement learning (RL) and associated dopaminergic processes have modulatory effects. For example, interval timing, which includes the judgment of relatively short intervals of time (ranging from milliseconds to…
Pierre Thodoroff, Audrey Durand, Joëlle Pineau, Doina Precup
Several applications of Reinforcement Learning suffer from instability due to high variance. This is especially prevalent in high dimensional domains. Regularization is a commonly used technique in machine learning to reduce variance, at the cost of introducing some bias. Most existing regularization techniques focus…
Margarida Sousa, Pawel Bujalski, Bruno F. Cruz, Kenway Louie + 2 more
Learning to predict rewards is a fundamental driver of adaptive behavior. Midbrain dopamine neurons (DANs) play a key role in such learning by signaling reward prediction errors (RPEs) that teach recipient circuits about expected rewards given current circumstances and actions. However, the algorithm that DANs are…
Authors not listed
Accurate prediction of chemical reaction yields remains essential for accelerating synthesis optimization, yet current machine learning models face critical limitations in capturing temporal dynamics, providing calibrated uncertainty estimates, and explicitly modeling reactant-to-product transformations. Here we…
Patrick M. Pilarski, Andrew Butcher, Elnaz Davoodi, Michael Bradley Johanson + 6 more
'Michael Bradley Johanson' 'Dylan J. A. Brenneis' 'Adam S. R. Parker' 'Leslie Acker' 'Matthew Botvinick' 'Joseph Modayil' 'Adam White'] Patrick M. Pilarski1,3,4, Andrew Butcher1 , Elnaz Davoodi1 , Michael Bradley Johanson1 , Dylan J. A. Brenneis1 , Adam S. R. Parker3,4, Leslie Acker1 , Matthew M. Botvinick2 , Joseph…
André Luzardo, Eduardo Alonso, Esther Mondragón
Computational models of classical conditioning have made significant contributions to the theoretic understanding of associative learning, yet they still struggle when the temporal aspects of conditioning are taken into account. Interval timing models have contributed a rich variety of time representations and provided…
Luca R. Bruder, Lisa Scharer, Jan Peters
Increased reactivity to addiction related cues (cue-reactivity) plays a critical role in the maintenance of addiction. Studies assessing cue-reactivity in gambling disorder often suffer from low ecological validity due to the usage of picture stimuli in a neutral lab environment. Here we describe a novel virtual…
David Mathar, Annika Wiebe, Deniz Tuzsus, Kilian Knauth + 1 more
Computational psychiatry focuses on identifying core cognitive processes that appear altered across a broad range of psychiatric disorders. Temporal discounting of future rewards and model-based control during reinforcement learning have proven as two promising candidates. Despite its trait-like stability, temporal…
Michael N. Hallquist, Alexandre Y. Dombrovski
Laboratory studies of value-based decision-making often involve choosing among a few discrete actions. Yet in natural environments, we encounter a multitude of options whose values may be unknown or poorly estimated. Given that our cognitive capacity is bounded, in complex environments, it becomes hard to solve the…