21 papers · ranked by Valyu relevance
Mohammad Ali Humayun, Junaid Shuja, Pg Emeroylariffion Abas, Syed Sajid Ullah
'Syed Sajid Ullah'] Social background profiling of speakers is heavily used in areas, such as, speech forensics, and tuning speech recognition for accuracy improvement. This article provides a survey of recent research in speaker background profiling in terms of accent classification and analyses the datasets, speech…
Philippe Albouy, Samuel A. Mehr, Roxane S. Hoyer, Jérémie Ginzburg + 1 more
Humans produce two primary forms of vocal communication: speaking and singing. What is the basis for these two categories? Is the distinction between them based primarily on culturally specific, learned features, or do consistent acoustical cues exist that reliably distinguish speech and song worldwide? Some studies…
Ying-Yee Kong, Ala Mullangi, Kostas Kokkinakis, Fan-Gang Zeng
Objective To investigate a set of acoustic features and classification methods for the classification of three groups of fricative consonants differing in place of articulation. Method A support vector machine (SVM) algorithm was used to classify the fricatives extracted from the TIMIT database in quiet and also in…
M. N. Renukadevi, T. M. Rajesh, S. G. Shaila, Praveen Kulkarni + 3 more
Speaker vocal delivery significantly impacts audience engagement in knowledge-sharing platforms like TED Talks. However, existing classification approaches face limitations in noise handling, feature extraction, and clustering robustness. The study investigates vocal behavior using a curated set of 5,000 TED Talk audio…
Charvi Vitthal, B.S. Shreeharsha, Kamini Sabu, Preeti Rao
Literacy assessment is an important activity for education administrators across the globe. Typically achieved in a school setting by testing a child's oral reading, it is intensive in human resources. While automatic speech recognition (ASR) is a potential solution to the problem, it tends to be computationally…
T. Ananthapadmanabha, K V Vijay Girish, A. G. Ramakrishnan
Detection of transitions between broad phonetic classes in a speech signal is an important problem which has applications such as landmark detection and segmentation. The proposed hierarchical method detects silence to non-silence transitions, high amplitude (mostly sonorants) to low amplitude (mostly…
Zeinab Mahmoudi, Saeed Rahati, Mohammad Mahdi Ghasemi, Vahid Asadpour + 2 more
'Vahid Asadpour' 'Hamid Tayarani' 'Mohsen Rajati'] Background Speech production and speech phonetic features gradually improve in children by obtaining audio feedback after cochlear implantation or using hearing aids. The aim of this study was to develop and evaluate automated classification of voice disorder in…
Mircea Petrache, Andrés Carvallo, Valentina Silva, Pablo Barceló + 1 more
How informative are preschoolers’ speech vocalizations? Preschoolers’ speech is often imprecise, highly variable and hard to interpret by humans and machines; consequently, its predictive value for later developmental outcomes remains quite underexplored. Here, we analyzed 6.595 brief vocalizations (0.5-5s) from 127…
Santiago-Omar Caballero-Morales
An approach for the recognition of emotions in speech is presented. The target language is Mexican Spanish, and for this purpose a speech database was created. The approach consists in the phoneme acoustic modelling of emotion-specific vowels. For this, a standard phoneme-based Automatic Speech Recognition (ASR) system…
Neema Mishra, Urmila Shrawankar, V. M. Thakare
In this age of information technology, information access in a convenient manner has gained importance. Since speech is a primary mode of communication among human beings, it is natural for people to expect to be able to carry out spoken dialogue with computer [1]. Speech recognition system permits ordinary people to…
Jamil Ahmad, Mustansar Fiaz, Soonil Kwon, Maleerat Sodanil + 2 more
'Sung Wook Baik'] Abstract— Gender recognition is an essential component of automatic speech recognition and interactive voice response systems. Determining gender of the speaker reduces the computational burden of such systems for any further processing. Typical methods for gender recognition from speech largely…
Stav Hertz, Benjamin Weiner, Nisim Perets, Michael London
Many complex motor behaviors can be decomposed into sequences of simple individual elements. Mouse ultrasonic vocalizations (USVs) are naturally divided into distinct syllables and thus are useful for studying the neural control of complex sequences production. However, little is known about the rules governing their…
Mahdieh Ghazvini, Seyedamiryousef Hosseini Goki, Sajad Hamzenejadi
— Pitch and formant frequencies are important features in speech processing applications. The period of the vocal cord's output for vowels is known as the pitch or the fundamental frequency, and formant frequencies are essentially resonance frequencies of the vocal tract. These features vary among different persons and…
Takahiro Fukumori, Taito Ishida, Yoichi Yamashita
The detection of shouted speech is crucial in audio surveillance and monitoring. Although it is desirable for a security system to be able to identify emergencies, existing corpora provide only a binary label (i.e., shouted or normal) for each speech sample, making it difficult to predict the shout intensity.…
Assel Davletcharova, Sherin Sugathan, Bibia Abraham, Alex Pappachen James
'Alex Pappachen James'] Recognizing emotion from speech has become one the active research themes in speech processing and in applications based on human-computer interaction. This paper conducts an experimental study on recognizing emotions from human speech. The emotions considered for the experiments include…
Herath Mudiyanselage Dhammike Piyumal Madhurajith Herath, Weraniyagoda Arachchilage Sahanaka Anuththara Weraniyagoda, Rajapakshage Thilina Madhushan Rajapaksha, Patikiri Arachchige Don Shehan Nilmantha Wijesekara + 3 more
'Weraniyagoda Arachchilage Sahanaka Anuththara Weraniyagoda' 'Rajapakshage Thilina Madhushan Rajapaksha' 'Patikiri Arachchige Don Shehan Nilmantha Wijesekara' 'Kalupahana Liyanage Kushan Sudheera' 'Peter Han Joo Chong' 'Leon Rothkrantz'] Aphasia is a type of speech disorder that can cause speech defects in a person.…
Hariharan Muthusamy, Kemal Polat, Sazali Yaacob, Haipeng Peng
In the recent years, many research works have been published using speech related features for speech emotion recognition, however, recent studies show that there is a strong correlation between emotional states and glottal features. In this work, Mel-frequency cepstralcoefficients (MFCCs), linear predictive cepstral…
Girwan Dhakal, Hongsheng He, Sharlene Newman, Yanyu Xiong
Speech acts shape early language development and social cognition, yet little is known about how late-talking (LT) children use them to achieve communicative goals. We compared LT and typically developing (TD) preschoolers (1;09–6;00) across nine dyadic English corpora, using a Conditional Random Field model to…
K Wierucka, D Murphy, SK Watson, N Falk + 9 more
Automated acoustic analysis is increasingly used in animal communication studies, and determining caller identity is a key element for many investigations. However, variability in feature extraction and classification methods limits the comparability of results across species and studies, constraining conclusions we…
Maya Inbar, Eitan Grossman, Ayelet N. Landau
Intonation units (IUs) are a universal building-block of human speech. They are found cross-linguistically and are tied to important language functions such as the pacing of information in discourse and swift turn-taking. We study the rate of IUs in 48 languages from every continent and from 27 distinct language…
Authors not listed
Acoustic measurements of batteries are known to be correlated to their state-of-charge, creating opportunities for state estimation that do not rely on electrical signals. State estimators are typically parametric models fitted from data, often from the broad toolbox of machine learning. Such models can be easily…