006.3
Machine Intelligence
Learning systems, from theory to frontier models.
Drawer contents
Filled from
Search threads
deep learning new architecture results
reinforcement learning agents
006.3
Learning systems, from theory to frontier models.
Drawer contents
Filled from
Search threads
deep learning new architecture results
reinforcement learning agents
Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen
Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, forcing them to acquire task-relevant knowledge through online interaction. Existing approaches obtain informative initializations through pre-collected datasets…
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu + 9 more
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser…
Zishan Xu, Zhiyuan Yao, Yuxin Chen, Yifu Guo + 8 more
Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that…
Xingyu Guo, Wei Chen, Linlin Yang, Baochang Zhang
Search agents extend large language models beyond static parametric memory by enabling them to acquire and use ex ternal evidence during multi-step reasoning. For knowledge intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant ev idence but also on using it…
Katrin Schmid, Iuri Frosio
Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning. In the specific case of video games, access to the game engine is required to get rewards for training (e.g. to collect rewards from the environment). Furthermore, the proper identification and weighting of the rewards…
Xuanyu Lei, Yiqi Zhu, Chenliang Li, Kaiming Liu + 5 more
Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with handcrafted verifiers. Though effective, both remain bottlenecked by externally specified tasks and supervision signals, limiting the scalability and diversity of agent…
Chengyang He, Tanishq Duhan, Gadiel Sznaier Camps, Fangyuan Wang + 5 more
We present PRIMAL3, an ultra-large-scale learning-based framework for multi-agent pathfinding (MAPF) that integrates reinforcement learning, topology-aware communication, LaCAM3-guided training, and PIBT-based action refinement. PRIMAL3 targets failures at topologically critical states, where agents must coordinate…
Patrick Krauss, Achim Schilling, Andreas Maier, Thomas Kinfe + 1 more
Deep Belief Networks (DBNs) learn hierarchical generative models without class supervision. Here, we ask whether this purely unsupervised process nevertheless organizes internal representations according to the unknown data classes. We analyze successive layers of DBNs trained on MNIST, Fashion-MNIST, and KMNIST using…
Hexiang Zhang, Mauro Antezza, Yi Zheng
We propose and analyze a programmable near-field radiative thermal network that emulates convolutional and recurrent operations through heat exchange. By integrating radiative thermal diodes and transistors composed of phase-change materials such as VO2 and GST, we construct two fundamental architectures: the Thermal…
Alexander Auras, Martin Burger, Samira Kabri, Michael Moeller + 1 more
Deep neural networks have shown great empirical success in the solution of a wide variety of ill-posed inverse problems in imaging. Yet, very few works have studied their behavior in the limit that turns the discretized ill-conditioned problems into truly ill-posed ones, i.e., for an increasing resolution of the…
Seyed Roozbeh Razavi Rohani, Khashayar Khajavi, Wesley Chung, Mandana Samiei + 1 more
Continual learning (CL) requires models to learn tasks sequentially, yet deep neural networks often suffer from plasticity loss and poor knowledge transfer, which can impede their long-term adaptability. Drawing high-level inspiration from global neuromodulatory mechanisms in the brain, we introduce Neuromodulation and…
Kai Li, Conggai Li, Sarah Ali Siddiqui, Syed Sohail Ahmed + 3 more
Agentic artificial intelligence (AI) is transforming Integrated Sensing and Communication (ISAC) from a function-oriented physical-layer technology into a goal-driven, closed-loop intelligent system, a paradigm we term AISAC. Existing work on learning-based sensing, resource allocation, reconfigurable intelligent…
Mahshad Rastegarmoghaddam, Davoud Nikkhouy, Shima Samadzadeh
Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control. We develop an integrated architecture in which the uncertainty estimate updates the obstacle geometry used by a control barrier…
Xiaofeng Wang, Kakam Chong, Shuai Xiao, DeXin Kong + 8 more
Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination and rapid adaptation. Current methods often apply uniform goal-based rewards to every utterance, overlooking the specificity of objectives at each dialogue turn and…
Weiwei Li, Junzhuo Liu, Tong Chu, Hengfu Yu + 1 more
GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each trajectory into prefix-action pairs: the agent predicts an action from the current screen and interaction history, while the subsequent observation is discarded. This removes the rationale of why an…
Vaishnav Vaidheeswaran, Dilith Jayakody, Biruk Ambaw, Jaswanth Kumar + 2 more
Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deployment requires reward models that are interpretable and robust to changing environments. Inverse reinforcement learning (IRL) provides a framework for recovering such rewards from…
Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari
Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents. However, most existing systems primarily optimize task success while giving limited consideration to execution efficiency under practical constraints such as executor capability and computational…
Hongjiang Wang, Weizhe Wang, Yingzheng Liu
Efficient solution of time-dependent parametric partial differential equations (PDEs) is central to computational science and engineering. Existing deep-learning-based accelerators span a physics-to-data spectrum, from physics-informed solvers with strong physical consistency but high computational cost to data-driven…
Songpan Gao, Yajie Zhang, Guanxing Chen, Jiayu Qian + 8 more
Deep learning models applied to medical image analysis suffer from severe catastrophic forgetting when continually adapting to new clinical tasks in dynamic environments. Mainstream incremental learning methods typically mitigate this by rehearsing raw historical images. However, this pixel-level rehearsal incurs…
Andrew L. Smith, Linxing Preston Jiang, Jason K. Eshraghian, Matthew S. Bull + 1 more
Hierarchical predictive coding proposes a compelling hypothesis of brain computation, suggesting that the cortex builds layered predictions to minimize surprise. Yet most models rely on error-coding neurons or generative modeling of unclear biological plausibility. Here, we examine a biologically plausible framework in…
Maryam Gholami Shiri, Eva Tuba, Sašo Džeroski, Tome Eftimov + 1 more
Benchmarking deep learning (DL) models for multi-label classification (MLC) of remote sensing images (RSI) typically yields rankings that do not generalize beyond the evaluated datasets. In this work, we move beyond rankings by employing functional analysis of variance (fANOVA) to systematically quantify the…
Zirui Chen, Shi Tang, Zhengchao Gao, Yongjia Su + 2 more
Although the state-of-the-art neural network model extraction attack in the hard-label setting by Carlini et al. at EUROCRYPT 2025 has polynomial-time complexity in theory, its dual-point clustering relies on singular value decomposition (SVD) with a time complexity of $\mathcal{O}(n^2 \cdot (d^{(k)})^3)$, resulting in…