16 papers · ranked by Valyu relevance
Marlies Gillis, Jill Kries, Jan Wouters, Laura Gwilliams + 1 more
This study investigates the neural dynamics of phoneme processing in 7-year-old children with and without dyslexia (25;9 ♂), using EEG recordings collected during continuous speech listening. By applying temporal generalization to phonetic descriptor decoding, we can disentangle whether potential phoneme processing…
Zhang, Yizi, He, Linyang + 20 more
Speech brain–computer interfaces (BCIs) aim to restore communication for people with paralysis by translating neural activity into text. Most systems use cascaded frameworks that decode phonemes before assembling sentences with an n-gram language model (LM), preventing joint optimization of all stages simultaneously.…
Oli Danyi Liu, Hao Tang, Naomi H. Feldman, Sharon Goldwater
Speech representations in the human brain do not simply mirror the instantaneous speech signal; rather, they display several properties that are hypothesized to facilitate the integration of speech sounds into words. In particular, neural encodings of speech maintain information that has dissipated from the acoustics…
Tommaso Boccato, Michal Olak, Matteo Ferrante
Brain-to-text systems have recently achieved impressive performance when trained on single-participant data, but remain limited by uninvestigated cross-subject generalization. We present the first neural-to-phoneme decoder trained jointly on the two largest intracortical speech datasets (Willett et al. 2023; Card et…
Anna Chrabaszcz, Kailee Lear, Corrine Durisko, Julie Fiez
This study investigated whether adults with long-standing reading difficulties (“poor readers”) can acquire new literacy skills in both familiar (English pseudowords) and novel (artificial orthography, AO) systems, and how phonological decoding deficits relate to orthographic learning. Poor readers (n = 17) and matched…
Jill Kries, Maaike Vandermosten, Laura Gwilliams
During successful language comprehension, speech sounds (phonemes) are encoded within a series of neural patterns that evolve over time. Here we tested whether these neural dynamics of speech encoding are altered for individuals with a language disorder. We recorded EEG responses from the human brains of 39 individuals…
Haodong Zhang, Wai Ting Siok, Nizhuan Wang, Wan-Young Chung
Imagined speech decoding has attracted growing interest in brain-computer interface (BCI) research, as it may enable language-related information to be recovered from non-overt neural activity. Current studies in this area are often treated as a single, unified research problem, despite substantial differences in…
Nghia Hieu Nguyen, Quan Ngoc Hoang, Long Hoang Huu Nguyen, Kiet Van Nguyen + 1 more
Most Automatic Speech Recognition (ASR) systems formulate transcription as a prediction problem over orthographic units such as characters, subwords, or words. Although effective, such representations do not explicitly reflect the phonetic structure of speech and often require large vocabularies to maintain adequate…
Julien Gadonneix, Mingfang Zhang, Jérémy Rapin, Linnea Evanson + 2 more
Speech production requires the rapid coordination of a complex hierarchy of linguistic units, transforming a semantic representation into a precise sequence of articulatory movements. To unravel the neural mechanisms underlying this feat, we leverage recordings from eight 3.2 x 3.2 mm 64-microelectrode arrays implanted…
Benyamin Abramovich Krasa, Erin M. Kunz, Foram Kamdar, Donald Avansino + 11 more
Speech requires precise serial ordering of words and phonemes into novel combinations. To accomplish this, the brain is believed to flexibly prepare utterances before producing them, even allowing pronunciation of never-before spoken words. To discover how neural populations achieve this, intracortical activity from…
Scott Crossley, Joon Suh Choi, Kenny Tang, Laurie Cutting
This study documents and assesses the Tool for Automatic Analysis of Decoding Ambiguity (TAADA). TAADA calculates measures related to decoding, including metrics for grapheme and phoneme counts, neighborhood effects, rhymes, and conditional probabilities for sound-spelling relationships. These measures are assessed in…
Lucas Zamora Vera, Jose A. Gonzalez-Lopez
State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent decoders, and how the output target (phonetic vs.\ character) interacts with that choice. On…
Quan Ngoc Hoang, Long Hoang Huu Nguyen, Nghia Hieu Nguyen, Kiet Van Nguyen + 1 more
Vietnamese exhibits substantial dialectal phonetic variation across Northern, Central, and Southern regions, where identical lexical items may be realized with markedly different pronunciations. Such variation poses challenges for automatic speech recognition (ASR) and remains difficult to model computationally due to…
Seung-Cheol Baek, Seung-Goo Kim, Burkhard Maess, Maren Grigutsch + 1 more
Prosody is a fundamental aspect of speech characterized by suprasegmental features such as pitch. Prosodic pitch contours are used to convey speakers’ intentions, for example, to make a statement or ask a question. Understanding these intentions requires abstracting continuous, variable pitch information into discrete…
Gavin M. Bidelman, Zara Eisenhut, Lucy Borowski, Rose Rizzi + 1 more
Speech perception requires that listeners classify sensory information into smaller groupings while also coping with noise that often corrupts the speech signal. The strength of categorization and speech-in-noise (SIN) abilities show stark individual differences. Some listeners perceive speech sounds in a gradient…
Abner Hernandez, Tomás Arias-Vergara, Daiqi Liu, Andreas Maier + 1 more
Phonological features provide a language-general and linguistically grounded representation of speech. We present PhonoQ-2.0, a multilingual frame-level phonological feature recognizer built on self-supervised speech models. The system directly predicts a structured 22-dimensional feature vector per frame encoding…