English
Related papers

Related papers: DualCodec: A Low-Frame-Rate, Semantically-Enhanced…

200 papers

Neural Audio Codecs (NACs) have become increasingly adopted in speech processing tasks due to their excellent rate-distortion performance and compatibility with Large Language Models (LLMs) as discrete feature representations for audio…

Sound · Computer Science 2025-09-15 Harry Julian , Rachel Beeson , Lohith Konathala , Johanna Ulin , Jiameng Gao

How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in…

We introduce Gull, a generative multifunctional audio codec. Gull is a general purpose neural audio compression and decompression model which can be applied to a wide range of tasks and applications such as real-time communication, audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 Yi Luo , Jianwei Yu , Hangting Chen , Rongzhi Gu , Chao Weng

Large Audio Language Models (LALMs) demonstrate impressive performance across diverse tasks, ranging from speech recognition to general audio understanding. However, their scalability is limited by the quadratic complexity of attention and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-27 Saurabhchand Bhati , Samuel Thomas , Hilde Kuehne , Rogerio Feris , James Glass

Noise reduction is an important part of modern hearing aids and is included in most commercially available devices. Deep learning-based state-of-the-art algorithms, however, either do not consider real-time and frequency resolution…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-29 Hendrik Schröter , Tobias Rosenkranz , Alberto N. Escalante B. , Marc Aubreville , Andreas Maier

Denoising language models (DLMs) have been proposed as a powerful alternative to traditional language models (LMs) for automatic speech recognition (ASR), motivated by their ability to use bidirectional context and adapt to a specific ASR…

Neural and Evolutionary Computing · Computer Science 2025-12-16 Dorian Koch , Albert Zeyer , Nick Rossenbach , Ralf Schlüter , Hermann Ney

Recent advances in neural audio codec-based speech generation (CoSG) models have produced remarkably realistic audio deepfakes. We refer to deepfake speech generated by CoSG systems as codec-based deepfake, or CodecFake. Although existing…

Sound · Computer Science 2025-08-05 Xuanjun Chen , I-Ming Lin , Lin Zhang , Jiawei Du , Haibin Wu , Hung-yi Lee , Jyh-Shing Roger Jang

We introduce a novel technique for creative audio resynthesis that operates by reworking the concept of granular synthesis at the latent vector level. Our approach creates a "granular codebook" by encoding a source audio corpus into latent…

Sound · Computer Science 2025-07-28 Nao Tokui , Tom Baker

Neural audio signal codecs have attracted significant attention in recent years. In essence, the impressive low bitrate achieved by such encoders is enabled by learning an abstract representation that captures the properties of encoded…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-06 Mhd Modar Halimeh , Matteo Torcoli , Philipp Grundhuber , Emanuël A. P. Habets

Audio tokenization bridges continuous waveforms and multi-track music language models. In dual-track modeling, tokens should preserve three properties at once: high-fidelity reconstruction, strong predictability under a language model, and…

Sound · Computer Science 2026-04-02 Rui Lin , Zhiyue Wu , Jiahe Le , Kangdi Wang , Weixiong Chen , Junyu Dai , Tao Jiang

With the rapid advancement of large language models (LLMs), discrete speech representations have become crucial for integrating speech into LLMs. Existing methods for speech representation discretization rely on a predefined codebook size…

Sound · Computer Science 2025-01-03 Linqin Wang , Yaping Liu , Zhengtao Yu , Shengxiang Gao , Cunli Mao , Yuxin Huang , Wenjun Wang , Ling Dong

Prevalent semantic speech tokenizers, designed to capture linguistic content, are surprisingly fragile. We find they are not robust to meaning-irrelevant acoustic perturbations; even at high Signal-to-Noise Ratios (SNRs) where speech is…

Computation and Language · Computer Science 2026-04-15 Yuhan Song , Linhao Zhang , Chuhan Wu , Aiwei Liu , Wei Jia , Houfeng Wang , Xiao Zhou

Deep learning-based models have greatly advanced the performance of speech enhancement (SE) systems. However, two problems remain unsolved, which are closely related to model generalizability to noisy conditions: (1) mismatched noisy…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-29 Cheng Yu , Ryandhimas E. Zezario , Syu-Siang Wang , Jonathan Sherman , Yi-Yen Hsieh , Xugang Lu , Hsin-Min Wang , Yu Tsao

Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for audio-and-text large language models (LLMs). Previous…

This paper presents a semantic-enhanced receiver framework for transmitting natural language sentences over noisy wireless channels using multiple short block codes. After ASCII encoding, the sentence is divided into segments, each…

Information Theory · Computer Science 2026-04-30 Jiafu Hao , Chentao Yue , Wanchun Liu , Branka Vucetic , Yonghui Li

Neural speech codecs have revolutionized speech coding, achieving higher compression while preserving audio fidelity. Beyond compression, they have emerged as tokenization strategies, enabling language modeling on speech and driving…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-02 Wei-Cheng Tseng , David Harwath

Autoregressive (AR) frameworks have recently achieved remarkable progress in zero-shot text-to-speech (TTS) by leveraging discrete speech tokens and large language model techniques. Despite their success, existing AR-based zero-shot TTS…

Sound · Computer Science 2025-10-14 Jingyuan Xing , Mingru Yang , Zhipeng Li , Xiaofen Xing , Xiangmin Xu

The voice mode of the Opus audio coder can compress wideband speech at bit rates ranging from 6 kb/s to 40 kb/s. However, Opus is at its core a waveform matching coder, and as the rate drops below 10 kb/s, quality degrades quickly. As the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Jan Skoglund , Jean-Marc Valin

Neural Audio Codecs (NACs) can reduce transmission overhead by performing compact compression and reconstruction, which also aim to bridge the gap between continuous and discrete signals. Existing NACs can be divided into two categories:…

Sound · Computer Science 2026-01-07 Zhisheng Zhang , Xiang Li , Yixuan Zhou , Jing Peng , Shengbo Cai , Guoyang Zeng , Zhiyong Wu

Although recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant. A generative…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-23 Yi-Chiao Wu , Dejan Marković , Steven Krenn , Israel D. Gebru , Alexander Richard
‹ Prev 1 8 9 10 Next ›