20 papers · ranked by Valyu relevance
Maoning Ge, Kento Ohtani, Yingjie Niu, Yuxiao Zhang + 3 more
Autonomous driving in complex real-world environments requires robust perception, reasoning, and physically feasible planning, which remain challenging for current end-to-end approaches. This paper introduces VLA-MP, a unified vision-language-action framework that integrates multimodal Bird’s-Eye View (BEV) perception…
Diana C. Dima, Sugitha Janarthanan, Jody C. Culham, Yalda Mohsenzadeh
Humans can recognize and communicate about many actions performed by others. How are actions organized in the mind, and is this organization shared across vision and language? We collected similarity judgments of human actions depicted through naturalistic videos and sentences, and tested four models of action…
Theodor Wulff, Federico Tavella, Rahul Singh Maharjan, Manith Adikari + 1 more
Achieving robot transparency is a critical step toward effective human-robot collaboration. To be transparent, a robot's natural language communication must be consistent with its actions and explicitly grounded in the task and environment. Existing hierarchical Vision-Language-Action (VLA) models can generate language…
Bin Ren, Diwei Shi, Enrico Meli
Highlights What are the main findings?1. A model of memory-gated filtering attention was proposed, which improved multi-head self-attention mechanism. 2. A cross-modal alignment perception during training was designed, which combined with a few-shot data collection strategy of key steps. What are the implications of…
Jiaqi Li, Guangming Wang, Shuntian Zheng, Minzhe Ni + 3 more
Temporal Action Localization (TAL) requires identifying both the boundaries and categories of actions in untrimmed videos. While visionlanguage models (VLMs) offer rich semantics to complement visual evidence, existing approaches tend to overemphasize linguistic priors at the expense of visual performance, leading to a…
Patrick Cavanagh
The descriptions of surfaces, objects, and events computed by visual processes are not solely for consumption in the visual system but are meant to be passed on to other brain centers. Clearly, the description of the visual scene cannot be sent in its entirety, like a picture or movie, to other centers, as that would…
Homagni Saha, Fateme Fotouhi, Qisai Liu, Soumik Sarkar
In this paper we propose a new framework-MoViLan (Modular Vision and Language) for execution of visually grounded natural language instructions for day to day indoor household tasks. While several data-driven, end-to-end learning frameworks have been proposed for targeted navigation tasks based on the vision and…
Massimo Bosetti, Shibingfeng Zhang, Benedetta Liberatori, Giacomo Zara + 2 more
'Giacomo Zara' 'Elisa Ricci' 'Paolo Rota'] Abstract. Vision-language models (VLMs) have demonstrated remarkable performance across various visual tasks, leveraging joint learning of visual and textual representations. While these models excel in zero-shot image tasks, their application to zero-shot video action…
Moritz F. Wurm, Alfonso Caramazza
We understand actions from both observation and written text, pointing to a common neural representation of action concepts. However, which parts of the brain encode action concepts independently of stimulus type, say language or visual observation, is an unresolved question in neuroscience. An overlap of activation…
Ozan Özdemir, Matthias Kerzel, Cornelius Weber, Jae Hee Lee + 1 more
'Stefan Wermter'] Abstract—Human infants learn language while interacting with their environment in which their caregivers may describe the objects and actions they perform. Similar to human infants, artificial agents can learn language while interacting with their environment. In this work, first, we present a neural…
Diana C. Dima, Jody C. Culham, Yalda Mohsenzadeh
Actions are the building blocks of our dynamic visual world, yet the neural computations supporting action perception are not well understood. How does perceptual and conceptual information unfold in the brain when we observe what others are doing? We collected EEG and fMRI data while participants viewed short videos…
Jisu Hwang, Incheol Kim, Miguel Arevalillo-Herráez
Due to the development of computer vision and natural language processing technologies in recent years, there has been a growing interest in multimodal intelligent tasks that require the ability to concurrently understand various forms of input data such as images and text. Vision-and-language navigation (VLN) require…
Matteo Ruggero Ronchi, Pietro Perona
Which common human actions and interactions are recognizable in monocular still images? Which involve objects and/or other people? How many is a person performing at a time? We address these questions by exploring the actions and interactions that are detectable in the images of the MS COCO dataset. We make two main…
Authors not listed
Molecular property prediction is a fundamental task in computational chemistry with critical applications in drug discovery and materials science. While recent works have explored Large Language Models (LLMs) for this task, they primarily rely on textual molecular representations such as SMILES/SELFIES, which can be…
Jane Han, Vassiki Chauhan, Rebecca Philip, Morgan K. Taylor + 5 more
We effortlessly extract behaviorally relevant information from dynamic visual input in order to understand the actions of others. In the current study, we develop and test different classes of models to better understand the neural representational geometries supporting action understanding. Using fMRI, we measured…
Joshua Shepherd
I argue that the neural realizers of experiences of trying (that is, experiences of directing effort towards the satisfaction of an intention) are not distinct from the neural realizers of actual trying (that is, actual effort directed towards the satisfaction of an intention). I then ask how experiences of trying…
Moritz F. Wurm, Seoyoung Lee
Higher-level action interpretation, such as inferring underlying intentions and predicting future actions, requires the integration of conceptual action information (e.g. "opening") with semantic knowledge about persons and objects (e.g. "my friend Anna", "pizza box"). However, how the neural systems for action and…
Ece Takmaz, Sandro Pezzelle, Raquel Fernández
the Variation in Human Signals during Visuo-Linguistic Processes Authors: ['Ece Takmaz' 'Sandro Pezzelle' 'Raquel Fernández'] There is an intricate relation between the properties of an image and how humans behave while describing the image. This behavior shows ample variation, as manifested in human signals such as…
David F. Marks
The Action Cycle Theory (ACT) is an enactive theory of the perception and a mental imagery system that is comprised of six modules: Schemata, Objects, Actions, Affect, Goals and Others’ Behavior. The evidence supporting these six connected modules is reviewed in light of research on mental imagery vividness. The six…
Authors not listed
In experimental chemistry, actions are adjusted based on what we see—such as dosing until dissolution, heating until melting, or stirring until mixing is complete. However, current self-driving labs (SDLs) do not monitor these visual cues. HeinSight 4.0 fills this gap by integrating computer vision into SDLs to enable…