Search · four archives
Search · four archives
24 papers · ranked by Valyu relevance
Yifan Zhong, Bai, Fengshuo, Shaofei Cai + 12 more
The remarkable advancements of vision and language foundation models in multimodal understanding, reasoning, and generation has sparked growing efforts to extend such intelligence to the physical world, fueling the flourishing of vision-language-action (VLA) models. Despite seemingly diverse approaches, we observe that…
Diana C. Dima, Sugitha Janarthanan, Jody C. Culham, Yalda Mohsenzadeh
Humans can recognize and communicate about many actions performed by others. How are actions organized in the mind, and is this organization shared across vision and language? We collected similarity judgments of human actions depicted through naturalistic videos and sentences, and tested four models of action…
Maoning Ge, Kento Ohtani, Yingjie Niu, Yuxiao Zhang + 3 more
Autonomous driving in complex real-world environments requires robust perception, reasoning, and physically feasible planning, which remain challenging for current end-to-end approaches. This paper introduces VLA-MP, a unified vision-language-action framework that integrates multimodal Bird’s-Eye View (BEV) perception…
Byoung Chul Ko, Yongmin Zhong, Bijan Shirinzadeh, Gaoge Hu + 1 more
Recent Vision-Language-Action (VLA) models have rapidly emerged as general-purpose robotic policies that integrate language understanding, visual perception, and robot control. However, prior studies and surveys have primarily emphasized backbone architectures, action decoders, training recipes, and benchmark…
Liu, Zhenyang, Gu, Yongchong + 8 more
Recent advancements in vision-language models (VLMs) for common-sense reasoning have led to the development of vision-language-action (VLA) models, enabling robots to perform generalized manipulation. Although existing autoregressive VLA methods design a specific architecture like dual-system to leverage large-scale…
Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao + 1 more
—Deep learning has demonstrated remarkable success across many domains, including computer vision, natural language processing, and reinforcement learning. Representative artificial neural networks in these fields span convolutional neural networks, Transformers, and deep Q-networks. Built upon unimodal neural…
Fuhao Li, Song, Wenxuan, Zhao + 6 more
Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are built upon vision-language models pretrained solely on 2D data, which lack accurate spatial awareness and hinder their ability to operate in the…
Bin Ren, Diwei Shi, Enrico Meli
Highlights What are the main findings?1. A model of memory-gated filtering attention was proposed, which improved multi-head self-attention mechanism. 2. A cross-modal alignment perception during training was designed, which combined with a few-shot data collection strategy of key steps. What are the implications of…
Zhihao Wang, Jianxiong Li, Jinliang Zheng, Wencong Zhang + 5 more
Vision-Language-Action (VLA) models have achieved notable success but often struggle with limited generalizations. To address this, integrating generalized Vision-Language Models (VLMs) as assistants to VLAs has emerged as a popular solution. However, current approaches often combine these models in rigid, sequential…
Dewen Zhang, Tahir Hussain, Wangpeng An, Hayaru Shouno + 1 more
Current vision-language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized vision-language instruction-following data. We introduce a method for generating such…
Xiaoyan Dai
1## Introduction Vision-Language Models (VLMs) have become a central paradigm in multimodal artificial intelligence. Large-scale image-text pretraining has enabled strong performance in image captioning, visual question answering, visual instruction following, and increasingly complex multimodal reasoning. These…
Jiaqi Shi, Xulong Zhang, Xiaoyang Qu, Jianzong Wang
Recent advances in Vision-Language-Action (VLA) models have shown promise for robot control, but their dependence on action supervision limits scalability and generalization. To address this challenge, we introduce CARE, a novel framework designed to train VLA models for robotic task execution. Unlike existing methods…
Jae Hee Lee, Yuan Yao, Ozan Özdemir, Mengdi Li + 3 more
'Zhiyuan Liu' 'Stefan Wermter'] A cognitive agent performing in the real world needs to learn relevant concepts about its environment (e.g., objects, color, and shapes) and react accordingly. In addition to learning the concepts, it needs to learn relations between the concepts, in particular spatial relations between…
Vinit Mehta, Charu Sharma, Karthick Thiyagarajan, Nader Jalili + 1 more
With the rapid advancement of artificial intelligence and robotics, the integration of Large Language Models (LLMs) with 3D vision is emerging as a transformative approach to enhancing robotic sensing technologies. This convergence enables machines to perceive, reason, and interact with complex environments through…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Daniel Fleury
From an early age, humans are challenged with evaluating rich environments full of socially and physically grounded concepts. For example, we might be spectating a rapidly unfolding tennis match, anticipating ball trajectories based on players’ body cues and goals. In another scenario, we may engage with long…
Diana C. Dima, Jody C. Culham, Yalda Mohsenzadeh
Actions are the building blocks of our dynamic visual world, yet the neural computations supporting action perception are not well understood. How does perceptual and conceptual information unfold in the brain when we observe what others are doing? We collected EEG and fMRI data while participants viewed short videos…
Jane Han, Vassiki Chauhan, Rebecca Philip, Morgan K. Taylor + 5 more
We effortlessly extract behaviorally relevant information from dynamic visual input in order to understand the actions of others. In the current study, we develop and test different classes of models to better understand the neural representational geometries supporting action understanding. Using fMRI, we measured…
Authors not listed
Molecular property prediction is a fundamental task in computational chemistry with critical applications in drug discovery and materials science. While recent works have explored Large Language Models (LLMs) for this task, they primarily rely on textual molecular representations such as SMILES/SELFIES, which can be…
Moritz F. Wurm, Seoyoung Lee
Higher-level action interpretation, such as inferring underlying intentions and predicting future actions, requires the integration of conceptual action information (e.g. "opening") with semantic knowledge about persons and objects (e.g. "my friend Anna", "pizza box"). However, how the neural systems for action and…
Heeseung Lee, Daeho Kim, Heyin Lee, Namyoung Gwak + 6 more
- 1. Computational Science Research Center, Korea Institute of Science and Technology, Seoul 02792, Republic of Korea - 2. Department of Materials Science and Engineering, Korea University, 145 Anam-ro, Seoul 02841, Republic of Korea - 3. Department of Chemical and Biological Engineering, Korea University, Seoul 02841…
Authors not listed
Natural language processing with the help of large language models such as ChatGPT has become ubiquitous in many software applications and allows users to interact even with complex hardware or software in an intuitive way. The recent concepts of Self-Driving Labs and Material Acceleration Platforms stand to benefit…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
W. Dupont, C. Papaxanthis, F. Lebon, C. Madden-Lombardi
Action reading is thought to engage motor simulations, yielding modulations in activity of motor-related cortical regions, and contributing to action language comprehension. To test these ideas, we measured 1) corticospinal excitability during action reading, and 2) reading comprehension ability, in individuals with…