22 papers · ranked by Valyu relevance
Amirhossein Khalilian-Gourtani, Chenqian Le, Faxin Zhou, Erika Jenson + 6 more
Speech is a defining human behavior, and this ability depends critically on speech motor cortex. While the ventral precentral and postcentral gyri are classically regarded as chiefly articulatory and somatosensory regions, a growing body of literature challenges this simplification. Most prior research, however, has…
Nicholas S. Card, Tyler Singer-Clark, Hamza Peracha, Carrina Iacobacci + 7 more
Brain-computer interfaces (BCIs) can provide naturalistic communication and digital access to people with severe paralysis by decoding neural activity associated with attempted speech and movement. Recent work has demonstrated highly accurate intracortical BCIs for speech and cursor control, but two critical…
Tommaso Boccato, Michal Olak, Matteo Ferrante
Brain-to-text systems have recently achieved impressive performance when trained on single-participant data, but remain limited by uninvestigated cross-subject generalization. We present the first neural-to-phoneme decoder trained jointly on the two largest intracortical speech datasets (Willett et al. 2023; Card et…
Chenyu Tang, Shuo Gao, Cong Li, Wentian Yi + 19 more
Wearable silent speech systems hold significant potential for restoring communication in patients with speech impairments. However, seamless, coherent speech remains elusive, and clinical efficacy is still unproven. Here, we present an AI-driven intelligent throat (IT) system that integrates throat muscle vibrations…
Mahmoud Keshavarzi, Brian C. J. Moore, Usha Goswami
Neural oscillations in the delta (0.5–4 Hz) and theta (4–8 Hz) bands play a key role in tracking the temporal structure of speech. According to Temporal Sampling (TS) theory, dyslexia arises from atypical entrainment of these low-frequency oscillations to speech during infancy and childhood, which is particularly…
Jiawei Li, Kaixuan Bian, Xiaotao Hao, Jinsong Wu + 2 more
Face-to-face communication relies on the seamless integration of visual and acoustic cues, yet the spatiotemporal principles governing how the human brain dynamically represents and combines these multisensory streams remain largely unresolved. To address this, we recorded high-density electrocorticography (ECoG) from…
Jill Kries, Maaike Vandermosten, Laura Gwilliams
During successful language comprehension, speech sounds (phonemes) are encoded within a series of neural patterns that evolve over time. Here we tested whether these neural dynamics of speech encoding are altered for individuals with a language disorder. We recorded EEG responses from the human brains of 39 individuals…
Aparna Srinivasan, Maitreyee Wairagkar, Carrina Iacobacci, Xianda Hou + 9 more
The ability to vary the mode and loudness of speech is an important part of the expressive range of human vocal communication. However, the encoding of these behaviors in the ventral precentral gyrus (vPCG) has not been studied at the resolution of neuronal firing rates. We investigated this in two participants who had…
Chen Feng, En Zhang, Yifei Jia, Zhoule Zhu + 3 more
Introduction Accurate and reliable detection of speech state transitions is a prerequisite for practical speech brain-computer interfaces (BCIs). While cortical language areas have been extensively studied, it remains unclear whether speech onset information is exclusively localized to these regions or distributed…
Oli Danyi Liu, Hao Tang, Naomi H. Feldman, Sharon Goldwater
Speech representations in the human brain do not simply mirror the instantaneous speech signal; rather, they display several properties that are hypothesized to facilitate the integration of speech sounds into words. In particular, neural encodings of speech maintain information that has dissipated from the acoustics…
Jinghan Yang, Haoran Jiang, Yanru Bai, Guangjian Ni + 1 more
Rapid advancements in artificial intelligence (AI) have enabled text-to-speech (TTS) systems to produce voices increasingly indistinguishable from humans, posing significant societal risks, particularly through potential misuse in fraud and deception. To address this concern, this study combined behavioral assessments…
Richard Csaky, Mats W.J. van Es, Oiwi Parker Jones, Mark Woolrich
Despite the prevalence of inner speech in everyday life, research on this has been limited, particularly when it comes to non-invasive methods. This preprint aims to fill this gap by using EEG and MEG to collect data from three different inner speech paradigms, and by conducting an initial decoding analysis.…
Esra Sümer-Arpak, Rajkumar Saini, Debashis Das Chakladar, Sanjeev Kumar Varun + 1 more
Inner speech (IS), or imagined speech without overt articulation, is a promising target for brain-computer interfaces (BCIs) aimed at restoring communication in individuals with severe speech impairments, such as locked-in syndrome. Foundation models (FMs), typically trained using self-supervised learning (SSL) on…
Lucas Zamora Vera, Jose A. Gonzalez-Lopez
State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent decoders, and how the output target (phonetic vs.\ character) interacts with that choice. On…
Guangting Mai, Emily Upton, Timothy D Griffiths, Alexander P Leff + 11 more
Comprehending connected speech is critical for human interaction and is vulnerable in post-stroke aphasia. Understanding the neural mechanisms underlying impaired speech listening is necessary for accurate and effective assessment and treatment. Neural speech tracking methods offer a window into naturalistic speech…
Marlies Gillis, Jill Kries, Jan Wouters, Laura Gwilliams + 1 more
This study investigates the neural dynamics of phoneme processing in 7-year-old children with and without dyslexia (25;9 ♂), using EEG recordings collected during continuous speech listening. By applying temporal generalization to phonetic descriptor decoding, we can disentangle whether potential phoneme processing…
Xiao-Hang Jiang, Yang Ai, Rui-Chen Zheng, Li-Rong Dai + 2 more
Neural speech codecs are key to speech transmission and storage, but most use uniform quantization across frames, allocating the same bitrate regardless of content and wasting bits. We propose VoCodec, a low-bitrate streamable neural speech codec with voicing-driven quantization that assigns higher bitrate to voiced…
Siyu Wang, Haitao Li, Zhu, Donglai
— Voice communication in bandwidth-constrained environments—maritime, satellite, and tactical networks—remains prohibitively expensive. Traditional codecs struggle below 1 kbps, while existing semantic approaches (STT-TTS) sacrifice prosody and speaker identity. We present STCTS, a generative semantic compression…
Hounsu Kim, Juhan Nam
Speaker-decoupled speech codecs can reduce bitrate by separating global speaker attributes from local content and prosody, while supporting voice conversion. Existing speaker-decoupled codecs face a trade-off: methods that explicitly suppress speaker leakage often rely on multi-stage or auxiliary training, whereas…
Tao Li, Wentong Ge, Zhichao Wang, Zihao Cui + 5 more
Codec-based language models (LMs) have revolutionized text-to-speech (TTS). However, standard codecs entangle timbre and prosody, which hinders independent control in continuation-based LMs. To tackle this challenge, we propose DisCo-Speech, a zeroshot controllable TTS framework featuring a disentangled speech codec…
Bhanu Teja Nellore, Sudarsana Reddy Kadiri, Rohit Kumar, Karan Nathwani + 1 more
It is well known that intelligibility of speech reduces in the presence of ambient noise. However, studies show that all sounds are not affected uniformly (or equally) and that vowels are more robust to noise than consonants. In this study, intelligibility of various consonants is assessed and analyzed in stationary…
Alex Gichamba, Moise Busogi
Low frame rates in neural audio codecs are attractive for autoregressive speech synthesis, where the generation cost scales linearly with the sequence length. Recent work has demonstrated that codecs can operate at 12.5 Hz and below, but the mechanisms underlying low frame rate degradation remain insufficiently…