15 papers · ranked by Valyu relevance
Wan, Zixiang, Zhang, Guochang + 4 more
Neural Audio Codecs (NACs) have gained growing attention in recent years as technologies for audio compression and audio representation in speech language models. While mainstream NACs typically require G-level computation and M-level parameters, the performance of lightweight and streaming NACs remains underexplored.…
Junyi Wang, Chi Zhang, Jing Qian, Haifeng Luo + 3 more
In bandwidth-constrained communication such as satellite and underwater channels, speech must often be transmitted at ultra-low bitrates where intelligibility is the primary objective. At such extreme compression levels, codecs trained with acoustic reconstruction losses tend to allocate bits to perceptual detail…
Junyi Wang, Chi Zhang, Jing Qian, Haifeng Luo + 3 more
In bandwidth-constrained communication such as satellite and underwater channels, speech must often be transmitted at ultra-low bitrates where intelligibility is the primary objective. At such extreme compression levels, codecs trained with acoustic reconstruction losses tend to allocate bits to perceptual detail…
Duong, Thien T., Springer, Jan P.
Perceptual quality of audio is the combination of aural accuracy and listener-perceived sound fidelity. It is how humans respond to the accuracy, intelligibility, and fidelity of aural media. Today this fidelity is also heavily influenced by the use of audio compression codecs for storing aural media in digital form.…
Hui-Peng Du, Yang Ai, Xiao-Hang Jiang, Yuan Tian + 1 more
Ultra-low-bitrate speech coding is pivotal for bandwidth-constrained communication and deep compression, yet maintaining naturalness and speaker identity at such extreme bit budgets remains challenging due to pronounced information loss and quantization instability. To this end, we propose FMelCodec, an…
Zhisheng Zhang, Xiang Li, Yi Zhou, Jing Peng + 3 more
—Neural Audio Codecs (NACs) can reduce transmission overhead by performing compact compression and reconstruction, which also aim to bridge the gap between continuous and discrete signals. Existing NACs can be divided into two categories: multi-codebook and single-codebook codecs. Multicodebook codecs face challenges…
Siyu Wang, Haitao Li, Zhu, Donglai
— Voice communication in bandwidth-constrained environments—maritime, satellite, and tactical networks—remains prohibitively expensive. Traditional codecs struggle below 1 kbps, while existing semantic approaches (STT-TTS) sacrifice prosody and speaker identity. We present STCTS, a generative semantic compression…
Xusheng Yang, Long Zhou, Wenfu Wang, Kai Hu + 5 more
We propose U-Codec, an Ultra low frame-rate neural speech Codec that achieves high-fidelity reconstruction and fast speech generation at an extremely low framerate of 5Hz (5 frames per second). Extreme compression at 5Hz typically leads to severe intelligibility and spectral detail loss, we introduce a…
Hao Ma, Rui Jing, Shansong Liu, Cheng Gong + 3 more
High-fidelity general audio compression at ultra-low bitrates is crucial for applications ranging from low-bandwidth communication to generative audio-language modeling. Traditional audio compression methods and contemporary neural codecs are fundamentally designed for waveform reconstruction. As a result, when…
Leyan Yang, Rui Hu, Yang Xu, Jing Lu
Recent advancements in end-to-end neural speech codecs enable compressing audio at extremely low bitrates while maintaining high-fidelity reconstruction. Meanwhile, low computational complexity and low latency are crucial for realtime communication. In this paper, we propose VoCodec, a speech codec model featuring a…
Sinclair Gurny, Ryan Quinn
Acoustic gunshot detection is a problem with applications across civilian public safety, military operations, and wildlife conservation, yet the field lacks a rigorous exploration of feature extraction techniques with a focus on generalization to realistic data. The mixed effectiveness of commercial gunshot detection…
Mingyu Zhao, Zijian Lin, Kun Wei, Zhiyong Wu
Conventional neural speech codecs suffer from severe intelligibility degradation at ultra-low bitrates, where the bottleneck transitions from acoustic distortion to semantic loss. To address this issue, this paper conducts a systematic investigation into the role and fundamental limits of integrating frozen semantic…
Ke-Han Lu, Keqi Deng, Ruchao Fan, Rui Zhao + 1 more
Speech large language models (Speech LLMs) typically encode speech into sequences far longer than text, creating a major efficiency bottleneck during autoregressive decoding. A common remedy is to compress the speech sequence at the adapter level to remove temporal redundancy before it enters the LLM; however, such…
Jun Xu, Zhengxue Cheng, Fengxi Zhang, Yuhan Liu + 2 more
Learning-based speech compression has achieved promising low-bitrate performance, but many neural speech codecs still describe quantized latents with preset-rate discrete symbols or apply entropy coding only after symbol generation. Such designs decouple representation learning from probability modeling, limiting their…
Sebastian, Rinku, O'Keefe Simon, Trefzer Martin
—Extracting features from the speech is the most critical process in Speech signal processing. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used features in the majority of the speaker and speech recognition applications as the filtering in this feature is similar to the filtering taking place in…