22 papers · ranked by Valyu relevance
Benjamin T. Files, Bosco S. Tjan, Jintao Jiang, Lynne E. Bernstein
From phonetic features to connected discourse, every level of psycholinguistic structure including prosody can be perceived through viewing the talking face. Yet a longstanding notion in the literature is that visual speech perceptual categories comprise groups of phonemes (referred to as visemes), such as /p, b, m/…
Authors not listed
Lip reading is used to understand or interpret speech without hearing it, a technique especially mastered by people with hearing difficulties. The ability to lip read enables a person with a hearing impairment to communicate with others and to engage in social activities, which otherwise would be difficult. Recent…
Patrick J. Karas, John F. Magnotti, Brian A. Metzger, Lin L. Zhu + 3 more
Vision provides a perceptual head start for speech perception because most speech is “mouth-leading”: visual information from the talker’s mouth is available before auditory information from the voice. However, some speech is “voice-leading” (auditory before visual). Consistent with a model in which vision modulates…
Aaron Nidiffer, Aisling O’Sullivan, Edmund C Lalor
In noisy environments, visible speech articulations improve listening comprehension. The benefit derives from several sources, including articulatory timing and shape. Recent research has shown that visual cortex encodes a categorical representation of articulatory features and that visual speech can benefit both…
Lynne E. Bernstein, Einat Liebenthal
This paper examines the questions, what levels of speech can be perceived visually, and how is visual speech represented by the brain? Review of the literature leads to the conclusions that every level of psycholinguistic speech structure (i.e., phonetic features, phonemes, syllables, words, and prosody) can be…
Samuel Evans, Cathy J. Price, Jörn Diedrichsen, Tae Twomey + 3 more
Reading is central to academic and vocational success. Some deaf children face reading challenges due to limited access to spoken or signed language. Robust phonological representations are key to reading development in hearing children. Spoken language phonology may be one of many contributors to reading development…
Maëva Michon, Gonzalo Boncompte, Vladimir López
The human brain generates predictions about future events. During face-to-face conversations, visemic information is used to predict upcoming auditory input. Recent studies suggest that the speech motor system plays a role in these cross-modal predictions, however, usually only audio-visual paradigms are employed. Here…
Aaron R Nidiffer, Cody Zhewei Cao, Aisling O’Sullivan, Edmund C Lalor
There is considerable debate over how visual speech is processed in the absence of sound and whether neural activity supporting lipreading occurs in visual brain areas. Surprisingly, much of this ambiguity stems from a lack of behaviorally grounded neurophysiological findings. To address this, we conducted an…
Ahmad B. Hassanat
Lip reading is used to understand or interpret speech without hearing it, a technique especially mastered by people with hearing difficulties. The ability to lip read enables a person with a hearing impairment to communicate with others and to engage in social activities, which otherwise would be difficult. Recent…
Adriana Fernandez-Lopez, Oriol Martínez, Federico M. Sukno
—Speech is the most used communication method between humans and it involves the perception of auditory and visual channels. Automatic speech recognition focuses on interpreting the audio signals, although the video can provide information that is complementary to the audio. Exploiting the visual information, however…
Adriana Fernandez-Lopez, Federico M. Sukno
Speech is the most common communication method between humans and involves the perception of both auditory and visual channels. Automatic speech recognition focuses on interpreting the audio signals, but it has been demonstrated that video can provide information that is complementary to the audio. Thus, the study of…
Prashant Bordea, Amarsinh Varpeb, Ramesh Manzac, Pravin Yannawara
Automatic Speech Recognition (ASR) by machine is an attractive research topic in signal processing domain and has attracted many researchers to contribute in this area. In recent year, there have been many advances in automatic speech reading system with the inclusion of audio and visual speech features to recognize…
Ahyeon Choi, Hayoon Kim, Mina Jo, Subeen Kim + 3 more
This review examines how visual information enhances speech perception in individuals with hearing loss, focusing on the impact of age, linguistic stimuli, and specific hearing loss factors on the effectiveness of audiovisual (AV) integration. While existing studies offer varied and sometimes conflicting findings…
Natalya Kaganovich, Jennifer Schumaker, Courtney Rowland
Background Visual speech cues influence different aspects of language acquisition. However, whether developmental language disorders may be associated with atypical processing of visual speech is unknown. In this study, we used behavioral and ERP measures to determine whether children with a history of SLI (H-SLI)…
Gilles Vannuscorps, Michael Andres, Sarah Carneiro, Elise Rombaux + 1 more
All it takes is a face to face conversation in a noisy environment to realize that viewing a speaker’s lip movements contributes to speech comprehension. Following the finding that brain areas that control speech production are also recruited during lip reading, the received explanation is that lipreading operates…
Souheil Fenghour, Daqing Chen, Kun Guo, Perry Xiao
The performance of automated lip reading using visemes as a classification schema has achieved less success compared with the use of ASCII characters and words largely due to the problem of different words sharing identical visemes. The Generative Pre-Training transformer is an effective autoregressive language model…
Jurgita Usinskiene, Michael Mouthon, Chrisovalandou Martins Gaytanidis, Agnes Toscanelli + 1 more
'Chrisovalandou Martins Gaytanidis' 'Agnes Toscanelli' 'Jean-Marie Annoni'] We propose a method of orthographic visualisation strategy in a poststroke severe aphasia person with dissociation between oral and written expression. fMRI results suggest that such strategy may induce the engagement of alternative nonlanguage…
Joan Orpella, Francesco Mantegna, M. Florencia Assaneo, David Poeppel
Speech imagery (the ability to generate internally quasi-perceptual experiences of speech events) is a fundamental ability tightly linked to important cognitive functions such as inner speech, phonological working memory, and predictive processing. Speech imagery is also considered an ideal tool to test theories of…
Authors not listed
Molecular property prediction is a fundamental task in computational chemistry with critical applications in drug discovery and materials science. While recent works have explored Large Language Models (LLMs) for this task, they primarily rely on textual molecular representations such as SMILES/SELFIES, which can be…
Authors not listed
For decades, molecular visualization software has been fundamental to education and research in chemistry, structural biology, and materials science. These tools have enabled the inspection of structures, dynamics, and interactions, yet their reliance on two-dimensional (2D) interfaces imposes persistent limitations.…
Authors not listed
The automatic generation of image captions in natural language is a critical and challenging task, particularly in the context of environmental monitoring and control. This paper presents a novel deep learning-driven image captioning system designed for real-time monitoring and predictive control of pollutant gas…
Authors not listed
This research presents a novel approach to obstacle detection during navigation using a combination of Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks. The primary objective is to generate accurate image captions that describe the content of images, which is crucial for applications such…