21 papers · ranked by Valyu relevance
Sandy Abdo, Bill Kapralos, Priyamvada Tripathi, KC Collins + 1 more
Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-driven audio generative models are rapidly growing in popularity and have the potential to transform…
Tiantian Feng, Dimitrios Dimitriadis, Shrikanth Narayanan
Recent advances in foundation models have enabled audiogenerative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models, the conventional approach to assessing the quality of the audio generation relies heavily on…
M. Huzaifah, L. Wyse
Sound modelling is the process of developing algorithms that generate sound under parametric control. There are a few distinct approaches that have been developed historically including modelling the physics of sound production and propagation, assembling signal generating and processing elements to capture acoustic…
Jonathan Morse, Naderi, Azadeh, Swen E. Gaudl + 3 more
Text-to-audio models are a type of generative model that produces audio output in response to a given textual prompt. Although level generators and the properties of the functional content that they create (e.g., playability) dominate most discourse in procedurally generated content (PCG), games that emotionally…
Fengrui Liu, Ruiyang Huang, Qijian Zheng, Yuanfang Wang + 1 more
Self-supervised learning advances audio representation for multimedia analysis. However, prevailing data-centric approaches rely on massive real-world corpora, increasing training costs, curation burdens, and privacy barriers. To address this, we present AudioPG, a procedural synthesis framework eliminating real audio…
Gašper Beguš
Training deep neural networks on well-understood dependencies in speech data can provide new insights into how they learn internal representations. This paper argues that acquisition of speech can be modeled as a dependency between random space and generated speech data in the Generative Adversarial Network…
Nikolaj Fišer, Miguel Ángel Martín-Pascual, Celia Andreu-Sánchez, Bruno Alejandro Mesz
'Bruno Alejandro Mesz'] Generative artificial intelligence (AI) has evolved rapidly, sparking debates about its impact on the visual and sonic arts. Despite its growing integration into creative industries, public opinion remains sceptical, viewing creativity as uniquely human. In music production, AI tools are…
Karl J. Friston, Noor Sajid, David Ricardo Quiroga-Martinez, Thomas Parr + 2 more
This paper introduces active listening, as a unified framework for synthesising and recognising speech. The notion of active listening inherits from active inference, which considers perception and action under one universal imperative: to maximise the evidence for our (generative) models of the world. First, we…
Gašper Beguš, Alan Zhou, T. Christina Zhao
Comparing artificial neural networks with outputs of neuroimaging techniques has recently seen substantial advances in (computer) vision and text-based language models. Here, we propose a framework to compare biological and artificial neural computations of spoken language representations and propose several new…
Andrey Anikin
Voice synthesis is a useful method for investigating the communicative role of different acoustic features. Although many text-to-speech systems are available, researchers of human nonverbal vocalizations and bioacousticians may profit from a dedicated simple tool for synthesizing and manipulating natural-sounding…
Tim Sainburg, Marvin Thielk, Timothy Q Gentner
Animals produce vocalizations that range in complexity from a single repeated call to hundreds of unique vocal elements patterned in sequences unfolding over hours. Characterizing complex vocalizations can require considerable effort and a deep intuition about each species’ vocal behavior. Even with a great deal of…
Shuji Komeiji, Kai Shigemi, Takumi Mitsuhashi, Yasushi Iimura + 5 more
Synthesizing speech from Electrocorticogram (ECoG) signals recorded during imagined speech remains a challenge due to the absence of synchronized audio signals for training. To address this, we propose a training framework that utilizes audio recorded during overt speech tasks as a surrogate ground truth for imagined…
Chitralekha Gupta, Purnima Kamath, Yize Wei, Zhuoyao Li + 2 more
'Suranga Nanayakkara' 'Lonce Wyse'] In this paper, we propose a data-driven approach to train a Generative Adversarial Network (GAN) conditioned on "soft-labels" distilled from the penultimate layer of an audio classifier trained on a target set of audio texture classes. We demonstrate that interpolation between such…
Hong Huang, Junfeng Man, Luyao Li, Rongke Zeng + 1 more
In this work, we focus on solving the problem of timbre transfer in audio samples. The goal is to transfer the source audio’s timbre from one instrument to another while retaining as much of the other musical elements as possible, including loudness, pitch, and melody. While image-to-image style transfer has been used…
Sara Popham, Dana Boebinger, Dan P. W. Ellis, Hideki Kawahara + 1 more
'Josh H. McDermott'] The “cocktail party problem” requires us to discern individual sound sources from mixtures of sources. The brain must use knowledge of natural sound regularities for this purpose. One much-discussed regularity is the tendency for frequencies to be harmonically related (integer multiples of a…
Nao Tokui
Since the introduction of deep learning, researchers have proposed content generation systems using deep learning and proved that they are competent to generate convincing content and artistic output, including music. However, one can argue that these deep learningbased systems imitate and reproduce the patterns…
Kat R. Agres, Adyasha Dash, Phoebe Chua
This work introduces a new music generation system, called AffectMachine-Classical, that is capable of generating affective Classic music in real-time. AffectMachine was designed to be incorporated into biofeedback systems (such as brain-computer-interfaces) to help users become aware of, and ultimately mediate, their…
Jongmin Ahn, Geun-Ho Park, Ho-Seuk Bae, Donghun Lee + 2 more
This study proposes a Variational AutoEncoder (VAE)-based Quality Controlled (QC) generative augmentation framework for North Atlantic right whale (NARW) upcall detection. Existing generative augmentation methods can generate synthetic samples; however, they do not provide a sample-level criterion for determining…
Artur Silva, Filipe Carvalho, Bruno F. Cruz
The design and characterization of a low-cost, open-source auditory delivery system to deliver high performance auditory stimuli is presented. The system includes a high-fidelity sound card and audio amplifier devices with low-latency and wide bandwidth targeted for behavioral neuroscience research. The…
Chad Williams, Daniel Weinhardt, Joshua Hewson, Martyna Beata Płomecka + 2 more
Electroencephalography (EEG) is a widely applied method for decoding neural activity, offering insights into cognitive function and driving advancements in neurotechnology. However, decoding EEG data remains challenging, as classification algorithms typically require large datasets that are expensive and time-consuming…
Authors not listed
Solving optimization problems, especially for nonlinear and constrained systems, is a challenge. Decades of specialized algorithms have been developed for general and special cases of root finding, minimization (including constraints), for parameter estimation, and mapping connected spaces. These approaches typically…