English
Related papers

Related papers: FunCodec: A Fundamental, Reproducible and Integrab…

200 papers

Diffusion speech enhancement on discrete audio codec features gain immense attention due to their improved speech component reconstruction capability. However, they usually suffer from high inference computational complexity due to multiple…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Yihui Fu , Tim Fingscheidt

Reconstructing natural images from functional magnetic resonance imaging (fMRI) data remains a core challenge in natural decoding due to the mismatch between the richness of visual stimuli and the noisy, low resolution nature of fMRI…

Image and Video Processing · Electrical Eng. & Systems 2025-09-03 Junliang Ye , Lei Wang , Md Zakir Hossain

This paper introduces FlowMAC, a novel neural audio codec for high-quality general audio compression at low bit rates based on conditional flow matching (CFM). FlowMAC jointly learns a mel spectrogram encoder, quantizer and decoder. At…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-08 Nicola Pia , Martin Strauss , Markus Multrus , Bernd Edler

In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot textbased speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Helin Wang , Meng Yu , Jiarui Hai , Chen Chen , Yuchen Hu , Rilin Chen , Najim Dehak , Dong Yu

Recently, the success of pre-training in text domain has been fully extended to vision, audio, and cross-modal scenarios. The proposed pre-training models of different modalities are showing a rising trend of homogeneity in their model…

Neural codec language models achieve impressive zero-shot Text-to-Speech (TTS) by fully imitating the acoustic characteristics of a short speech prompt, including timbre, prosody, and paralinguistic information. However, such holistic…

Sound · Computer Science 2026-01-21 Hanchen Pei , Shujie Liu , Yanqing Liu , Jianwei Yu , Yuanhang Qian , Gongping Huang , Sheng Zhao , Yan Lu

Neural audio compression has emerged as a promising technology for efficiently representing speech, music, and general audio. However, existing methods suffer from significant performance degradation at limited bitrates, where the available…

Sound · Computer Science 2026-05-08 Jin Wang , Wenbin Jiang , Xiangbo Wang , Yubo You , Sheng Fang

Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-19 Edresson Casanova , Ryan Langman , Paarth Neekhara , Shehzeen Hussain , Jason Li , Subhankar Ghosh , Ante Jukić , Sang-gil Lee

Streaming speech enhancement is a crucial task for real-time applications such as online meetings, smart home appliances, and hearing aids. Deep neural network-based approaches achieve exceptional performance while demanding substantial…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-29 Sunghwan Ahn , Jinmo Han , Beom Jun Woo , Nam Soo Kim

Neural speech codecs are essential for advancing text-to-speech (TTS) systems. With the recent success of large language models in text generation, developing high-quality speech tokenizers has become increasingly important. This paper…

Sound · Computer Science 2025-06-02 Peijie Chen , Wenhao Guan , Kaidi Wang , Weijie Wu , Hukai Huang , Qingyang Hong , Lin Li

Recent attempts at source tracing for codec-based deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using…

Sound · Computer Science 2025-08-19 Xuanjun Chen , I-Ming Lin , Lin Zhang , Haibin Wu , Hung-yi Lee , Jyh-Shing Roger Jang

The fundamental frequency (F0) contour of speech is a key aspect to represent speech prosody that finds use in speech and spoken language analysis such as voice conversion and speech synthesis as well as speaker and language identification.…

Audio and Speech Processing · Electrical Eng. & Systems 2018-05-09 Akihiro Kato , Tomi Kinnunen

Recent advancements in joint speech-text models show great potential for seamless voice interactions. However, existing models face critical challenges: temporal resolution mismatch between speech tokens (25Hz) and text tokens (~3Hz)…

Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-27 Xiaofei Wang , Manthan Thakker , Zhuo Chen , Naoyuki Kanda , Sefik Emre Eskimez , Sanyuan Chen , Min Tang , Shujie Liu , Jinyu Li , Takuya Yoshioka

Most neural vocoders are limited to one type: either GAN or diffusion-based. While state-of-the-art models like Vocos and WaveNeXt use powerful ConvNeXt-based generators, they have only been used in GAN frameworks and have limited…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Wangzixi Zhou , Takuma Okamoto , Yamato Ohtani , Sakriani Sakti , Hisashi Kawai

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge distillation objectives:…

Decoding continuous speech from intracortical recordings is a central challenge for brain-computer interfaces (BCIs), with transformative potential for individuals with conditions that impair their ability to speak. While recent…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Tommaso Boccato , Michal Olak , Matteo Ferrante

Although speech recognition algorithms have developed quickly in recent years, achieving high transcription accuracy across diverse audio formats and acoustic environments remains a major challenge. This work explores how incorporating…

Sound · Computer Science 2025-03-31 Aniket Abhishek Soni

High-fidelity neural audio codecs in Text-to-speech (TTS) aim to compress speech signals into discrete representations for faithful reconstruction. However, prior approaches faced challenges in effectively disentangling acoustic and…

Sound · Computer Science 2025-09-23 Ruonan Zhang , Xiaoyang Hao , Yichen Han , Junjie Cao , Yue Liu , Kai Zhang

Neural audio codecs have significantly advanced audio compression by efficiently converting continuous audio signals into discrete tokens. These codecs preserve high-quality sound and enable sophisticated sound generation through generative…

Sound · Computer Science 2025-02-12 Xiaoyu Bie , Xubo Liu , Gaël Richard