13 papers · ranked by Valyu relevance
Jakob Sponholz, Andreas Weilinghoff, Juliane Schopf
In qualitative research, data transcription is often labor-intensive and time-consuming. To expedite this process, a workflow utilizing artificial intelligence (AI) was developed. This workflow not only enhances transcription speed but also addresses the issue of AI-generated transcripts often lacking compatibility…
Kyra Wang, Dorien Herremans
Paralanguage Authors: ['Kyra Wang' 'Dorien Herremans'] Abstract—Laughing, sighing, stuttering, and other forms of paralanguage do not contribute any direct lexical meaning to speech, but they provide crucial propositional context that aids semantic and pragmatic processes such as irony. It is thus important for…
František Kmječ, Ondřej Bojar
Many meetings require creating a meeting summary to keep everyone up to date. Creating minutes of sufficient quality is however very cognitively demanding. Although we currently possess capable models for both audio speech recognition (ASR) and summarization, their fully automatic use is still problematic. ASR models…
MOSI. AI, Donghua Yu, Zhengyuan Lin, Chen Yang + 21 more
Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meeting transcription. Existing SATS systems rarely adopt an end-to-end formulation and are further constrained by limited context windows, weak…
Laurin Wagner, Mario Zusag, Bernhard Thallinger
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that…
Sunwoo Ha, Chaehun Lim, R. Jordan Crouser, Alvitta Ottley
Analysis and Exploration Authors: Sunwoo Ha, Chaehun Lim, R. Jordan Crouser, Alvitta Ottley Title: ConFides: A Visual Analytics Solution for Automated Speech Recognition Analysis and Exploration Authors: Sunwoo Ha, Chaehun Lim, R. Jordan Crouser, Alvitta Ottley Content: # 4.1.1 AWS Transcribe Output AWS Transcribe…
Mark D. Humphries, Lianne C. Leddy, Quinn Downton, Meredith Legace + 3 more
Handwritten Historical Documents Authors: ['Mark D. Humphries' 'Lianne C. Leddy' 'Quinn Downton' 'Meredith Legace' 'John McConnell' 'Iain Murray' 'Elizabeth Spence'] This study demonstrates that Large Language Models (LLMs) can transcribe historical handwritten documents with significantly higher accuracy than…
Max Bain, Jaesung Huh, Tengda Han, Andrew Zisserman
With the availability of large-scale web datasets, weaklysupervised and unsupervised training methods have demonstrated impressive performance on a multitude of speech processing tasks; including speech recognition , speaker recognition , speech separation , and keyword spotting . Whisper utilises this rich source of…
Vivek Senthil, Zhiqiang Tao, Ernest Fokoué
Modern policing faces a "visibility paradox" where law enforcement agencies possess petabytes of Body-Worn Camera (BWC) footage that remains largely unutilized for accountability or systemic review due to the prohibitive labor costs of manual transcription. This research presents a framework for adapting the OpenAI…
Cristina Aggazzotti, Nicholas Andrews, Elizabeth A. T. Smith
Authorship verification is the task of determining if two distinct writing samples share the same author and is typically concerned with the attribution of written text. In this paper, we explore the attribution of transcribed speech, which poses novel challenges. The main challenge is that many stylistic features…
Ashwin Rao
Videos are increasingly being used for e-learning, and transcripts are vital to enhance the learning experience. The costs and delays of generating transcripts can be alleviated by automatic speech recognition (ASR) systems. In this article, we quantify the transcripts generated by whisper for 25 educational videos and…
Minghan Wang, Yu-Xia Wang, Thuy-Trang Vu, Ehsan Shareghi + 1 more
Multimodal ASR Authors: ['Minghan Wang' 'Yu-Xia Wang' 'Thuy-Trang Vu' 'Ehsan Shareghi' 'Gholamreza Haffari'] Recent advancements in multimodal large language models (MLLMs) have made significant progress in integrating information across various modalities, yet real-world applications in educational and scientific…
Vladimir Bataev, Subhankar Ghosh, Vitaly Lavrukhin, Jason Li
—This work introduces TTS-Transducer – a novel architecture for text-to-speech, leveraging the strengths of audio codec models and neural transducers. Transducers, renowned for their superior quality and robustness in speech recognition, are employed to learn monotonic alignments and allow for avoiding using explicit…