Search · four archives
Search · four archives
25 papers · ranked by Valyu relevance
Abhijit Sen, Sonali Panda, Mahima Arya, Subhajit Patra + 2 more
Abhijit Sen , 1, ∗ Sonali Panda , 2, † Mahima Arya , 1, ‡ Subhajit Patra , 3, § Zizhan Zheng , 4, ¶ and Denys I. Bondar 1, ∗∗ 1 Department of Physics and Engineering Physics, Tulane University, New Orleans, LA 70118, USA 2 Department of Physics, Indian Institute of Technology, Dhanbad, India 3 Department of Electrical…
Lex Weaver, Jonathan Baxter
TD() with function approximation has proved empirically successful for some complex reinforcement learning problems. For linear approximation, TD() has been shown to minimise the squared error between the approximate value of each state and the true value. However, as far as policy is concerned, it is error in the…
Biru B. Dudhabhate, Kauê M. Costa
Dopamine signaling has become closely associated with reward prediction errors (RPEs)-the difference between expected and experienced value. Although not without controversy, the dopamine RPE hypothesis is one of the most influential ideas in neuroscience. This review briefly summarizes its origins, empirical…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Ghoshana Bista
Deep reinforcement learning has evolved from classical dynamic programming, temporal-difference learning, and tabular control into a broad framework for sequential decision-making under uncertainty. This book provides a structured introduction to that evolution, emphasizing not only how reinforcement learning…
Vasos Arnaoutis, Eric Lutters, Bojana Rosić
In this paper, we present a generalized temporal-difference (TD) reinforcement learning framework based on the theory of conditional expectations. The value and action-value (Q-value) functions are treated as uncertain quantities, and their estimation is formulated as a stochastic inference problem. Unlike classical…
Pranav Mahajan, Ben Seymour
The seminal reward prediction error account of dopamine has been highly successful, but faces several key challenges. Most notable are the difficulty of learning multiple rewards simultaneously, inefficient on-policy learning, and accounting for the heterogeneous striatal responses observed across and within striatal…
Jinyoung Jang, Juan C. Flores, Karen Zito, Randall C. O’Reilly
A major outstanding question in neuroscience is whether the neocortex uses the same powerful learning algorithm as current AI models: error backpropagation. One way this could be accomplished is as a function of the temporal derivative (i.e., differences in neural activity states over time), which can closely…
Adithya Gungi, Pradyumna Sepúlveda Delgado, Ines F. Aitsahalia, Marta Blanco-Pozo + 1 more
Flexible, goal-directed behavior depends on learning predictive relationships, yet how reward shapes learned transition structure remains incompletely understood. Here we introduce the Sparse Cognitive Graph, a reinforcement-learning framework in which a continuously updated transition representation is sparsified into…
Yifan Gao, Robert Wilson, Galit Karpov, Travis E. Baker
How does the brain learn to predict rewards? According to temporal difference (TD) learning theory, reward prediction errors (RPEs) should shift from the time of outcome delivery to earlier predictive cues as stimulus-action-outcome associations are learned. The reward positivity, an electrophysiological signal…
Pranav Mahajan, Ben Seymour
The seminal reward prediction error theory of dopamine function faces several key challenges. Most notable is the difficulty learning multiple rewards simultaneously, inefficient on-policy learning, and accounting for heterogeneous striatal responses in the tail of the striatum. We propose a normative framework, based…
Jinyoung Jang, Juan C. Flores, Karen Zito, Randall C. O'Reilly
A major outstanding question in neuroscience is whether the neocortex uses the same powerful learning algorithm as current AI models: error backpropagation. One way this could be accomplished is as a function of the temporal derivative (i.e., differences in neural activity states over time), which can closely…
Jessica Passlack, Andrew F. MacAskill, Alejandro Tabas
The ability to use context to flexibly adjust our decision-making is vital for navigating a complex world. To do this, the brain must both use environmental features and behavioural outcomes to distinguish between different, often hidden contexts; and also learn how to use these inferred contexts to guide behaviour.…
Charlotte Collingwood, Francesca Greenstreet, Marcus Stephenson-Jones, Rafal Bogacz
Action-selection is determined by a combination of goal-directed and habitual processes. Habits are defined as the reward-independent, stimulus-response relationships which form when an action is regularly executed in the same context, regardless of outcome. An influential computational model proposes that habit…
Sean R. Maulhardt, Alec Solway, Caroline J. Charpentier, Alireza Soltani
When receiving a reward after a sequence of multiple events, how do we determine which event caused the reward? This problem, known as temporal credit assignment, can be difficult for humans to solve given the temporal uncertainty in the environment. Research to date has attempted to isolate dimensions of delay and…
Brett Daley
Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a fundamental challenge. For 30 years, practitioners have been limited to a binary choice: eliminate the bias at the cost of severely truncated eligibility traces (Watkins' Q($λ$)), or…
Authors not listed
Accurate prediction of chemical reaction yields remains essential for accelerating synthesis optimization, yet current machine learning models face critical limitations in capturing temporal dynamics, providing calibrated uncertainty estimates, and explicitly modeling reactant-to-product transformations. Here we…
Perkins, Daniel, Escobar, Oscar J. + 2 more
We present a detailed study of Deep Q-Networks in finite environments, emphasizing the impact of epsilon-greedy exploration schedules and prioritized experience replay. Through systematic experimentation, we evaluate how variations in epsilon decay schedules affect learning efficiency, convergence behavior, and reward…
Takayuki Tsurumi, Kenji Morita
In learning goal-directed behavior, state representation is important for adapting to the environment and achieving goals. A predictive state representation called successive representation (SR) has recently attracted attention as a candidate for state representation in animal brains, especially in the hippocampus. The…
Carlos Antolín, Joan Falcó-Roget, Luis Serrano-Fernández, Néstor Parga
Decision-making depends on coordinated computations distributed across dorsal and ventral circuits, often described as actor–critic systems. We examined how this division of labor gives rise to value- and decision-related representations by training a bimodular recurrent network on an economic choice task and…
Dabal Pedamonti, Samia Mohinta, Martin V. Dimitrov, Hugo Malagon-Vina + 2 more
Mastering navigation in environments with limited visibility is crucial for survival. Although the hippocampus has been associated with goal-oriented navigation, its role in real-world behaviour remains unclear. To investigate this, we combined deep reinforcement learning (RL) modelling with behavioural and neural data…
Authors not listed
Molecular dynamics (MD) is a powerful tool for exploring the behavior of atomistic systems, but its reliance on sequential numerical integration limits simulation efficiency. We present MDtrajNet-1, a foundational AI model that directly generates MD trajectories across chemical space, bypassing force calculations and…
Authors not listed
Quantitative Structure Activity Relationship (QSAR) remains an effective tool for early-stage chemical modelling and virtual screening in drug design. The advancements in this field are led by two core paradigms, 1) descriptor engineering, where complex fixed-length vectors of compounds are generated and conventional…
Asim H. Gazi, Yongyi Guo, Daiqi Gao, Ziping Xu + 2 more
Reinforcement learning (RL) has achieved remarkable success in real-world decision-making across diverse domains, including gaming, robotics, online advertising, public health, and natural language processing. Despite these advances, a substantial gap remains between RL research and its deployment in many practical…
Authors not listed
The integration of machine learning methods is transforming many areas of research by, for instance, accelerating molecular dynamics simulations and enabling improved prediction and optimization of chemical reactions. However, despite this progress, the adoption of data-driven approaches in atomic layer deposition…