English
Related papers

Related papers: MagiCodec: Simple Masked Gaussian-Injected Codec f…

200 papers

Audio codecs are a critical component of modern speech generation systems. This paper introduces a low-bitrate, multi-scale residual codec that encodes speech into four distinct streams: semantic, timbre, prosody, and residual. This…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-25 Jingyu Li , Guangyan Zhang , Zhen Ye , Yiwen Guo

Generating realistic drum audio directly from symbolic representations is a challenging task at the intersection of music perception and machine learning. We propose a system that transforms an expressive drum grid, a time-aligned MIDI…

Diffusion-based audio and music generation models commonly perform generation by constructing an image representation of audio (e.g., a mel-spectrogram) and then convert it to audio using a phase reconstruction model or vocoder. Typical…

Sound · Computer Science 2024-10-08 Ge Zhu , Juan-Pablo Caceres , Zhiyao Duan , Nicholas J. Bryan

In this paper, we proposed AI-based audio coding using MFCC features in an adversarial setting. We combined a conventional encoder with an adversarial learning decoder to better reconstruct the original waveform. Since GAN gives implicit…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-24 Mohammad Reza Hasanabadi

In this paper, we propose a personalized neural speech codec, envisioning that personalization can reduce the model complexity or improve perceptual speech quality. Despite the common usage of speech codecs where only a single talker is…

Sound · Computer Science 2024-04-02 Inseon Jang , Haici Yang , Wootaek Lim , Seungkwon Beack , Minje Kim

Although recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant. A generative…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-23 Yi-Chiao Wu , Dejan Marković , Steven Krenn , Israel D. Gebru , Alexander Richard

Neural audio codecs, leveraging quantization algorithms, have significantly impacted various speech/audio tasks. While high-fidelity reconstruction is paramount for human perception, audio coding for machines (ACoM) prioritizes efficient…

Sound · Computer Science 2025-08-06 Anastasia Kuznetsova , Inseon Jang , Wootaek Lim , Minje Kim

Generative speech enhancement (GSE) models show great promise in producing high-quality clean speech from noisy inputs, enabling applications such as curating noisy text-to-speech (TTS) datasets into high-quality ones. However, GSE models…

Sound · Computer Science 2026-01-21 Kazuki Yamauchi , Masato Murata , Shogo Seki

A new achievable rate region is given for the Gaussian cognitive many-to-one interference channel. The proposed novel coding scheme is based on the compute-and-forward approach with lattice codes. Using the idea of decoding sums of…

Information Theory · Computer Science 2016-11-18 Jingge Zhu , Michael Gastpar

Audio denoising is critical in signal processing, enhancing intelligibility and fidelity for applications like restoring musical recordings. This paper presents a proof-of-concept for adapting a state-of-the-art neural audio codec, the…

Sound · Computer Science 2025-11-04 Daniel Jimon , Mircea Vaida , Adriana Stan

Generative image codecs aim to optimize perceptual quality, producing realistic and detailed reconstructions. However, they often overlook a key property of human vision: our tendency to focus on particular aspects of a visual scene (e.g.,…

Image and Video Processing · Electrical Eng. & Systems 2026-04-02 Lucas Relic , Roberto Azevedo , Yang Zhang , Stephan Mandt , Markus Gross , Christopher Schroers

This paper introduces a novel neural network-based speech coding system that can process noisy speech effectively. The proposed source-aware neural audio coding (SANAC) system harmonizes a deep autoencoder-based source separation model and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-11 Haici Yang , Kai Zhen , Seungkwon Beack , Minje Kim

Ultra-low-bitrate speech coding is pivotal for bandwidth-constrained communication and deep compression, yet maintaining naturalness and speaker identity at such extreme bit budgets remains challenging due to pronounced information loss and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Hui-Peng Du , Yang Ai , Xiao-Hang Jiang , Yuan Tian , Zhen-Hua Ling

Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebooks discretizing images into tokens for autoregressive modeling. Existing codebook-based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Yanzhe Chen , Huasong Zhong , Yan Li , Zhenheng Yang

With the proliferation of Large Language Model (LLM) based deepfake audio, there is an urgent need for effective detection methods. Previous deepfake audio generation methods typically involve a multi-step generation process, with the final…

For the additive Gaussian noise channel with average codeword power constraint, sparse superposition codes and adaptive successive decoding is developed. Codewords are linear combinations of subsets of vectors, with the message indexed by…

Information Theory · Computer Science 2010-06-22 Andrew R Barron , Antony Joseph

Recent advances in large language models (LLMs) have demonstrated impressive capabilities in code-related tasks, such as code generation and automated program repair. Despite their promising performance, most existing approaches for code…

Software Engineering · Computer Science 2025-09-03 Yicong Zhao , Shisong Chen , Jiacheng Zhang , Zhixu Li

Perceptual video compression leverages generative priors to reconstruct realistic textures and motions at low bitrates. However, existing perceptual codecs often lack native support for variable bitrate and progressive delivery, and their…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Daowen Li , Ruixiao Dong , Ying Chen , Kai Li , Ding Ding , Li Li

VQ-based image generation typically follows a two-stage pipeline: a tokenizer encodes images into discrete tokens, and a generative model learns their dependencies for reconstruction. However, improved tokenization in the first stage does…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Bin Wu , Mengqi Huang , Weinan Jia , Zhendong Mao

The multi-codebook speech codec enables the application of large language models (LLM) in TTS but bottlenecks efficiency and robustness due to multi-sequence prediction. To avoid this obstacle, we propose Single-Codec, a single-codebook…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Hanzhao Li , Liumeng Xue , Haohan Guo , Xinfa Zhu , Yuanjun Lv , Lei Xie , Yunlin Chen , Hao Yin , Zhifei Li