21 papers · ranked by Valyu relevance
Sultan Mahmud, Nahid Hasan, Kelsey Mankel, Mohammed Yeasin + 1 more
Categorical perception (CP) reflects the human auditory system’s ability to map continuous acoustic signals onto discrete categories. Understanding the relationship between neural dynamics and perceptual decisions is central to speech–language processing. Here, we implemented a data-driven approach using Bayesian…
Mohammad Ali Humayun, Junaid Shuja, Pg Emeroylariffion Abas, Syed Sajid Ullah
'Syed Sajid Ullah'] Social background profiling of speakers is heavily used in areas, such as, speech forensics, and tuning speech recognition for accuracy improvement. This article provides a survey of recent research in speaker background profiling in terms of accent classification and analyses the datasets, speech…
Philippe Albouy, Samuel A. Mehr, Roxane S. Hoyer, Jérémie Ginzburg + 1 more
Humans produce two primary forms of vocal communication: speaking and singing. What is the basis for these two categories? Is the distinction between them based primarily on culturally specific, learned features, or do consistent acoustical cues exist that reliably distinguish speech and song worldwide? Some studies…
Samson Akinpelu, Serestina Viriri
Speech emotion classification (SEC) has gained the utmost height and occupied a conspicuous position within the research community in recent times. Its vital role in Human-Computer Interaction (HCI) and affective computing cannot be overemphasized. Many primitive algorithmic solutions and deep neural network (DNN)…
Foteini Simistira Liwicki, Vibha Gupta, Rajkumar Saini, Kanjar De + 1 more
This study focuses on the automatic decoding of inner speech using noninvasive methods, such as electroencephalography (EEG)). While inner speech has been a research topic in philosophy and psychology for half a century, recent attempts have been made to decode nonvoiced spoken words by using various brain-computer…
Qiyue Wang, Yan Fu, Baiyu Shao, Le Chang + 3 more
'Yun Ling'] Parkinson’s disease (PD) is a neurodegenerative disorder that negatively affects millions of people. Early detection is of vital importance. As recent researches showed dysarthria level provides good indicators to the computer-assisted diagnosis and remote monitoring of patients at the early stages. It is…
Mahdieh Ghazvini, Seyedamiryousef Hosseini Goki, Sajad Hamzenejadi
— Pitch and formant frequencies are important features in speech processing applications. The period of the vocal cord's output for vowels is known as the pitch or the fundamental frequency, and formant frequencies are essentially resonance frequencies of the vocal tract. These features vary among different persons and…
Arnab Kumar Roy, Hemant Kumar Kathania, Paban Sapkota, Sudarsana Reddy Kadiri + 1 more
Dysarthric speech severity classification is challenging due to speaker variability, class imbalance, and limited datasets. This study introduces DSSCNet, a deep learning model that employs transfer learning and multi-corpus learning to enhance speaker-independent classification. By pre-training on one dysarthric…
M. N. Renukadevi, T. M. Rajesh, S. G. Shaila, Praveen Kulkarni + 3 more
Speaker vocal delivery significantly impacts audience engagement in knowledge-sharing platforms like TED Talks. However, existing classification approaches face limitations in noise handling, feature extraction, and clustering robustness. The study investigates vocal behavior using a curated set of 5,000 TED Talk audio…
G. A. Belokrylov, А. Коренев, B. Lodonova, A. Novokhrestov
The article describes an attempt to apply an ensemble of binary classifiers to solve the problem of speech assessment in medicine. A dataset was compiled based on quantitative and expert assessments of syllable pronunciation quality. Quantitative assessments of 7 selected metrics were used as features: dynamic time…
Mircea Petrache, Andrés Carvallo, Valentina Silva, Pablo Barceló + 1 more
How informative are preschoolers’ speech vocalizations? Preschoolers’ speech is often imprecise, highly variable and hard to interpret by humans and machines; consequently, its predictive value for later developmental outcomes remains quite underexplored. Here, we analyzed 6.595 brief vocalizations (0.5-5s) from 127…
Guidong Bao, Mengchen Lin, Xiaoqian Sang, Yangcan Hou + 2 more
'Yunfeng Wu'] This article proposes a novel semi-supervised competitive learning (SSCL) algorithm for vocal pattern classifications in Parkinson’s disease (PD). The acoustic parameters of voice records were grouped into the families of jitter, shimmer, harmonic-to-noise, frequency, and nonlinear measures, respectively.…
Yao-Ming Kuo, Shanq-Jang Ruan, Yu-Chin Chen, Ya‐Wen Tu
This article describes a system for analyzing acoustic data to assist in the diagnosis and classification of children's speech sound disorders (SSDs) using a computer. The analysis concentrated on identifying and categorizing four distinct types of Chinese SSDs. The study collected and generated a speech corpus…
M. Zakaria Kurdi
Early and accurate diagnosis of Alzheimer’s Disease (AD) is critical for effective intervention. While previous studies have explored speech-based biomarkers for AD, this paper presents the first systematic investigation of acoustic, prosodic, and phonological speech features for detecting this neurocognitive disorder.…
Takahiro Fukumori, Taito Ishida, Yoichi Yamashita
The detection of shouted speech is crucial in audio surveillance and monitoring. Although it is desirable for a security system to be able to identify emergencies, existing corpora provide only a binary label (i.e., shouted or normal) for each speech sample, making it difficult to predict the shout intensity.…
Kevin Huang, Anfeng Xu, Xuan Shi, Thanathai Lertpetchpun + 4 more
We present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English, Arabic, Mandarin and Cantonese, Tibetan, Indic languages, Thai, Spanish…
Herath Mudiyanselage Dhammike Piyumal Madhurajith Herath, Weraniyagoda Arachchilage Sahanaka Anuththara Weraniyagoda, Rajapakshage Thilina Madhushan Rajapaksha, Patikiri Arachchige Don Shehan Nilmantha Wijesekara + 3 more
'Weraniyagoda Arachchilage Sahanaka Anuththara Weraniyagoda' 'Rajapakshage Thilina Madhushan Rajapaksha' 'Patikiri Arachchige Don Shehan Nilmantha Wijesekara' 'Kalupahana Liyanage Kushan Sudheera' 'Peter Han Joo Chong' 'Leon Rothkrantz'] Aphasia is a type of speech disorder that can cause speech defects in a person.…
Benjamin Benti, Patrick JO Miller, Heike Vester, Florencia Noriega + 1 more
We present here an unsupervised procedure for the classification of graded animal vocalisations based on Mel frequency cepstral coefficients and fuzzy clustering. Cepstral coefficients compress information about the distribution of energy across the frequency spectrum into a reduced number of variables and are…
Akmalbek Bobomirzaevich Abdusalomov, Furkat Safarov, Mekhriddin Rakhimov, Boburkhon Turaev + 2 more
'Mekhriddin Rakhimov' 'Boburkhon Turaev' 'Taeg Keun Whangbo' 'Paolo Bellavista'] Speech recognition refers to the capability of software or hardware to receive a speech signal, identify the speaker’s features in the speech signal, and recognize the speaker thereafter. In general, the speech recognition process involves…
Aayush Kumar Sharma, Vineet Bhavikatti, Amogh Nidawani, Siddappaji + 2 more
'P Sanath' 'Dr Geetishree Mishra'] Abstract— In this research paper, we delve into the topics of Speech Diarization and Automatic Speech Recognition (ASR). Speech diarization involves the separation of individual speakers within an audio stream. By employing the ASR transcript, the diarization process aims to segregate…
Authors not listed
Acoustic measurements of batteries are known to be correlated to their state-of-charge, creating opportunities for state estimation that do not rely on electrical signals. State estimators are typically parametric models fitted from data, often from the broad toolbox of machine learning. Such models can be easily…