English
Related papers

Related papers: Gull: A Generative Multifunctional Audio Codec

200 papers

In this paper, we propose and investigate the use of neural audio codec language models for the automatic generation of sample-based musical instruments based on text or reference audio prompts. Our approach extends a generative audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-23 Shahan Nercessian , Johannes Imort , Ninon Devis , Frederik Blang

Granular sound synthesis is a popular audio generation technique based on rearranging sequences of small waveform windows. In order to control the synthesis, all grains in a given corpus are analyzed through a set of acoustic descriptors.…

Sound · Computer Science 2021-07-06 Adrien Bitton , Philippe Esling , Tatsuya Harada

Perceptual quality of audio is the combination of aural accuracy and listener-perceived sound fidelity. It is how humans respond to the accuracy, intelligibility, and fidelity of aural media. Today this fidelity is also heavily influenced…

Sound · Computer Science 2026-03-12 Thien T. Duong , Jan P. Springer

The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a comprehensive understanding across diverse audio types, such as…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Ziqian Wang , Xianjun Xia , Xinfa Zhu , Lei Xie

Built upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and auto-regressive audio generation. However, existing models face substantial challenges in perceptual quality and signal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Zhikang Niu , Sanyuan Chen , Long Zhou , Ziyang Ma , Xie Chen , Shujie Liu

Discrete speech tokenization is a fundamental component in speech codecs. However, in large-scale speech-to-speech systems, the complexity of parallel streams from multiple quantizers and the computational cost of high-time-dimensional…

Sound · Computer Science 2025-07-28 Rongkun Xue , Yazhe Niu , Shuai Hu , Zixin Yin , Yongqiang Yao , Jing Yang

While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. This gap persists largely because current benchmarks, limited by data annotations and evaluation…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-12 Yadong Niu , Tianzi Wang , Heinrich Dinkel , Xingwei Sun , Jiahao Zhou , Gang Li , Jizhong Liu , Xunying Liu , Junbo Zhang , Jian Luan

The proliferation of deep neural networks has spawned the rapid development of acoustic echo cancellation and noise suppression, and plenty of prior arts have been proposed, which yield promising performance. Nevertheless, they rarely…

Sound · Computer Science 2025-01-27 Zhihang Sun , Andong Li , Rilin Chen , Hao Zhang , Meng Yu , Yi Zhou , Dong Yu

In this paper we propose a novel model for unconditional audio generation based on generating one audio sample at a time. We show that our model, which profits from combining memory-less modules, namely autoregressive multilayer…

While recent machine learning research has revealed connections between deep generative models such as VAEs and rate-distortion losses used in learned compression, most of this work has focused on images. In a similar spirit, we view…

Image and Video Processing · Electrical Eng. & Systems 2024-10-28 Ruihan Yang , Yibo Yang , Joseph Marino , Stephan Mandt

Audio denoising is critical in signal processing, enhancing intelligibility and fidelity for applications like restoring musical recordings. This paper presents a proof-of-concept for adapting a state-of-the-art neural audio codec, the…

Sound · Computer Science 2025-11-04 Daniel Jimon , Mircea Vaida , Adriana Stan

Neural speech codecs aim to compress input signals into minimal bits while maintaining content quality in a low-latency manner. However, existing neural codecs often trade model complexity for reconstruction performance. These codecs…

Sound · Computer Science 2024-10-04 Yuzhe Gu , Enmao Diao

State-of-the-art non-autoregressive text-to-speech (TTS) models based on FastSpeech 2 can efficiently synthesise high-fidelity and natural speech. For expressive speech datasets however, we observe characteristic audio distortions. We…

Sound · Computer Science 2024-09-19 Fabian Kögel , Bac Nguyen , Fabien Cardinaux

Neural speech codecs have gained great attention for their outstanding reconstruction with discrete token representations. It is a crucial component in generative tasks such as speech coding and large language models (LLM). However, most…

Sound · Computer Science 2025-07-01 Youqiang Zheng , Weiping Tu , Yueteng Kang , Jie Chen , Yike Zhang , Li Xiao , Yuhong Yang , Long Ma

Recent advancements in audio generation have been significantly propelled by the capabilities of Large Language Models (LLMs). The existing research on audio LLM has primarily focused on enhancing the architecture and scale of audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-28 Zhen Ye , Peiwen Sun , Jiahe Lei , Hongzhan Lin , Xu Tan , Zheqi Dai , Qiuqiang Kong , Jianyi Chen , Jiahao Pan , Qifeng Liu , Yike Guo , Wei Xue

Retrieval-Augmented Generation (RAG) has gained significant attention in recent years for its potential to enhance natural language understanding and generation by combining large-scale retrieval systems with generative models. RAG…

Computation and Language · Computer Science 2025-03-18 Mingyue Cheng , Yucong Luo , Jie Ouyang , Qi Liu , Huijie Liu , Li Li , Shuo Yu , Bohou Zhang , Jiawei Cao , Jie Ma , Daoyu Wang , Enhong Chen

We consider the problem of simultaneous reduction of acoustic echo, reverberation and noise. In real scenarios, these distortion sources may occur simultaneously and reducing them implies combining the corresponding distortion-specific…

Sound · Computer Science 2020-07-28 Guillaume Carbajal , Romain Serizel , Emmanuel Vincent , Eric Humbert

Despite the recent progress on neural network architectures for speech separation, the balance between the model size, model complexity and model performance is still an important and challenging problem for the deployment of such models to…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-18 Yi Luo , Cong Han , Nima Mesgarani

We introduce SRC-gAudio, a novel audio generation model designed to facilitate text-to-audio generation across a wide range of sampling rates within a single model architecture. SRC-gAudio incorporates the sampling rate as part of the…

Sound · Computer Science 2024-10-10 Chenxing Li , Manjie Xu , Dong Yu

The Animation-based Generative Codec (AGC) is an emerging paradigm for talking-face video compression. However, deploying its intricate decoder on resource and power-constrained edge devices presents challenges due to numerous parameters,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Rui Wan , Qi Zheng , Ruoyu Zhang , Bu Chen , Jiaming Liu , Min Li , Minge Jing , Jinjia Zhou , Yibo Fan