English
Related papers

Related papers: FlowMAC: Conditional Flow Matching for Audio Codin…

200 papers

This paper proposes a method that enhances the compression performance of the current model under development for the upcoming MPEG standard on Feature Coding for Machines (FCM). This standard aims at providing inter-operable compressed…

Image and Video Processing · Electrical Eng. & Systems 2025-12-15 Juan Merlos , Fabien Racapé , Hyomin Choi , Mateen Ulhaq , Hari Kalva

Scalability and efficiency are desired in neural speech codecs, which supports a wide range of bitrates for applications on various devices. We propose a collaborative quantization (CQ) scheme to jointly learn the codebook of LPC…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-14 Kai Zhen , Mi Suk Lee , Jongmo Sung , Seungkwon Beack , Minje Kim

In voice conversion (VC) applications, diffusion and flow-matching models have exhibited exceptional speech quality and speaker similarity performances. However, they are limited by slow conversion owing to their iterative inference.…

Sound · Computer Science 2026-02-23 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Yuto Kondo

We introduce Matcha-TTS, a new encoder-decoder architecture for speedy TTS acoustic modelling, trained using optimal-transport conditional flow matching (OT-CFM). This yields an ODE-based decoder capable of high output quality in fewer…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-11 Shivam Mehta , Ruibo Tu , Jonas Beskow , Éva Székely , Gustav Eje Henter

In this paper, we presents a low-complexity deep learning frameworks for acoustic scene classification (ASC). The proposed framework can be separated into three main steps: Front-end spectrogram extraction, back-end classification, and late…

Sound · Computer Science 2021-06-17 Lam Pham , Hieu Tang , Anahid Jalali , Alexander Schindler , Ross King

Error Correction Codes (ECC) are fundamental to reliable digital communication, yet designing neural decoders that are both accurate and computationally efficient remains challenging. Recent denoising diffusion decoders achieve…

Machine Learning · Computer Science 2026-02-18 Haoyu Lei , Chin Wa Lau , Kaiwen Zhou , Nian Guo , Farzan Farnia

The Large Language models (LLMs) have demonstrated supreme capabilities in text understanding and generation, but cannot be directly applied to cross-modal tasks without fine-tuning. This paper proposes a cross-modal in-context learning…

Sound · Computer Science 2024-06-17 Dongchao Yang , Haohan Guo , Yuanyuan Wang , Rongjie Huang , Xiang Li , Xu Tan , Xixin Wu , Helen Meng

As consumer devices become increasingly intelligent and interconnected, efficient data transfer solutions for machine tasks have become essential. This paper presents an overview of the latest Feature Coding for Machines (FCM) standard,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Md Eimran Hossain Eimon , Juan Merlos , Ashan Perera , Hari Kalva , Velibor Adzic , Borko Furht

Speech signal processing is a cornerstone of modern communication technologies, tasked with improving the clarity and comprehensibility of audio data in noisy environments. The primary challenge in this field is the effective separation and…

Sound · Computer Science 2025-10-07 Abdulhady Abas Abdullah , Aram Mahmood Ahmed , Tarik Rashid , Hadi Veisi

Neural audio/speech coding has recently demonstrated its capability to deliver high quality at much lower bitrates than traditional methods. However, existing neural audio/speech codecs employ either acoustic features or learned blind…

Sound · Computer Science 2025-10-16 Xue Jiang , Xiulian Peng , Huaying Xue , Yuan Zhang , Yan Lu

Recent advances in video compression have seen significant coding performance improvements with the development of new standards and learning-based video codecs. However, most of these works focus on application scenarios that allow a…

Multimedia · Computer Science 2025-02-18 Siyue Teng , Yuxuan Jiang , Ge Gao , Fan Zhang , Thomas Davis , Zoe Liu , David Bull

Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video frames. Previous studies…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Shentong Mo , Yibing Song

Music codecs are a vital aspect of audio codec research, and ultra low-bitrate compression holds significant importance for music transmission and generation. Due to the complexity of music backgrounds and the richness of vocals, solely…

Sound · Computer Science 2025-07-14 Yaoxun Xu , Hangting Chen , Jianwei Yu , Wei Tan , Rongzhi Gu , Shun Lei , Zhiwei Lin , Zhiyong Wu

Neural audio codecs have recently gained traction for their ability to compress high-fidelity audio and provide discrete tokens for generative modeling. However, leading approaches often rely on resource-intensive models and complex…

Sound · Computer Science 2025-08-18 Linwei Zhai , Han Ding , Cui Zhao , fei wang , Ge Wang , Wang Zhi , Wei Xi

Neural audio codecs have recently enabled high-fidelity reconstruction at high compression rates, especially for speech. However, speech and non-speech audio exhibit fundamentally different spectral characteristics: speech energy…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-11 Haoran Wang , Jiatong Shi , Jinchuan Tian , Bohan Li , Kai Yu , Shinji Watanabe

Neural audio codecs (NACs) provide compact representations that can be leveraged in many downstream applications, in particular large language models. Yet most NACs encode mixtures of multiple sources in an entangled manner, which may…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-21 Ryo Aihara , Yoshiki Masuyama , Francesco Paissan , François G. Germain , Gordon Wichern , Jonathan Le Roux

Traditional low bit-rate speech coding approach only handles narrowband speech at 8kHz, which limits further improvements in speech quality. Motivated by recent successful exploration of deep learning methods for image and speech…

Sound · Computer Science 2021-02-05 Gang Min , Xiongwei Zhang , Xia Zou , Xiangyang Liu

This paper introduces a novel framework for end-to-end learned video coding. Image compression is generalized through conditional coding to exploit information from reference frames, allowing to process intra and inter frames with the same…

Image and Video Processing · Electrical Eng. & Systems 2021-04-29 Théo Ladune , Pierrick Philippe , Wassim Hamidouche , Lu Zhang , Olivier Déforges

Automated audio captioning (AAC) is an audio-to-text task to describe audio contents in natural language. Recently, the advancements in large language models (LLMs), with improvements in training approaches for audio encoders, have opened…

Sound · Computer Science 2024-06-26 Jizhong Liu , Gang Li , Junbo Zhang , Heinrich Dinkel , Yongqing Wang , Zhiyong Yan , Yujun Wang , Bin Wang

Generative models have shown robust performance on speech enhancement and restoration tasks, but most prior approaches operate offline with high latency, making them unsuitable for streaming applications. In this work, we investigate the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-21 Tsun-An Hsieh , Sebastian Braun
‹ Prev 1 4 5 6 7 8 10 Next ›