English
Related papers

Related papers: OmniCodec: Low Frame Rate Universal Audio Codec wi…

200 papers

We introduce Gull, a generative multifunctional audio codec. Gull is a general purpose neural audio compression and decompression model which can be applied to a wide range of tasks and applications such as real-time communication, audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 Yi Luo , Jianwei Yu , Hangting Chen , Rongzhi Gu , Chao Weng

Noise reduction is an important part of modern hearing aids and is included in most commercially available devices. Deep learning-based state-of-the-art algorithms, however, either do not consider real-time and frequency resolution…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-29 Hendrik Schröter , Tobias Rosenkranz , Alberto N. Escalante B. , Marc Aubreville , Andreas Maier

Recent neural audio compression models often rely on residual vector quantization for high-fidelity coding, but using a fixed number of per-frame codebooks is suboptimal for the wide variability of audio content-especially for signals that…

Sound · Computer Science 2026-05-08 Xiangbo Wang , Wenbin Jiang , Jin Wang , Yubo You , Sheng Fang , Fei Wen

We introduce ReALLM, a novel approach for compression and memory-efficient adaptation of pre-trained language models that encompasses most of the post-training quantization and fine-tuning methods for a budget of <4 bits. Pre-trained…

Machine Learning · Computer Science 2024-05-24 Louis Leconte , Lisa Bedin , Van Minh Nguyen , Eric Moulines

With the proliferation of Large Language Model (LLM) based deepfake audio, there is an urgent need for effective detection methods. Previous deepfake audio generation methods typically involve a multi-step generation process, with the final…

Speech tokenization enables discrete representation and facilitates speech language modeling. However, existing neural codecs capture low-level acoustic features, overlooking the semantic and contextual cues inherent to human speech. While…

Neural audio codecs, neural networks which compress a waveform into discrete tokens, play a crucial role in the recent development of audio generative models. State-of-the-art codecs rely on the end-to-end training of an autoencoder and a…

Sound · Computer Science 2025-03-26 Zineb Lahrichi , Gaëtan Hadjeres , Gael Richard , Geoffroy Peeters

Recent advancements in neural audio codecs have not only enabled superior audio compression but also enhanced speech synthesis techniques. Researchers are now exploring their potential as universal acoustic feature extractors for a broader…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-21 Wei-Cheng Tseng , David Harwath

Modern multimodal large language models (MLLMs) generate fluent responses from interleaved text, image, audio, and video inputs. However, identifying which input sources support each generated statement remains an open challenge. Existing…

Computation and Language · Computer Science 2026-04-16 Qianqi Yan , Yichen Guo , Ching-Chen Kuo , Shan Jiang , Hang Yin , Yang Zhao , Xin Eric Wang

Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio,…

Multimedia · Computer Science 2026-05-15 Jianghan Chao , Jianzhang Gao , Wenhui Tan , Yuchong Sun , Ruihua Song , Liyun Ru

Fine-grained perception of multimodal information is critical for advancing human-AI interaction. With recent progress in audio-visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel,…

Computation and Language · Computer Science 2026-03-17 Ziyang Ma , Ruiyang Xu , Zhenghao Xing , Yunfei Chu , Yuxuan Wang , Jinzheng He , Jin Xu , Pheng-Ann Heng , Kai Yu , Junyang Lin , Eng Siong Chng , Xie Chen

Recent advancements in multimodal large language models (MLLMs) have aimed to integrate and interpret data across diverse modalities. However, the capacity of these models to concurrently process and reason about multiple modalities remains…

Recent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis. However, effectively mapping the complex, multidimensional attributes of speech into discrete tokens remains…

State-of-the-art text-to-video generation models such as Sora 2 and Veo 3 can now produce high-fidelity videos with synchronized audio directly from a textual prompt, marking a new milestone in multi-modal generation. However, evaluating…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Susan Liang , Chao Huang , Filippos Bellos , Yolo Yunlong Tang , Qianxiang Shen , Jing Bi , Luchuan Song , Zeliang Zhang , Jason Corso , Chenliang Xu

Neural audio/speech coding has recently demonstrated its capability to deliver high quality at much lower bitrates than traditional methods. However, existing neural audio/speech codecs employ either acoustic features or learned blind…

Sound · Computer Science 2025-10-16 Xue Jiang , Xiulian Peng , Huaying Xue , Yuan Zhang , Yan Lu

Omni-modal Large Language Models (Omni-LLMs) have demonstrated strong capabilities in audio-video understanding tasks. However, their reliance on long multimodal token sequences leads to substantial computational overhead. Despite this…

Computation and Language · Computer Science 2026-05-14 Yue Ding , Yiyan Ji , Jungang Li , Xuyang Liu , Xinlong Chen , Junfei Wu , Bozhou Li , Bohan Zeng , Yang Shi , Yushuo Guan , Yuanxing Zhang , Jiaheng Liu , Qiang Liu , Pengfei Wan , Liang Wang

Neural speech codecs have recently emerged as a focal point in the fields of speech compression and generation. Despite this progress, achieving high-quality speech reconstruction under low-bitrate scenarios remains a significant challenge.…

Sound · Computer Science 2024-11-22 Yu Pan , Xiang Zhang , Yuguang Yang , Jixun Yao , Yanni Hu , Jianhao Ye , Hongbin Zhou , Lei Ma , Jianjun Zhao

Recent advances in large language models (LLMs) have demonstrated impressive capabilities in code-related tasks, such as code generation and automated program repair. Despite their promising performance, most existing approaches for code…

Software Engineering · Computer Science 2025-09-03 Yicong Zhao , Shisong Chen , Jiacheng Zhang , Zhixu Li

Neural Audio Codecs, initially designed as a compression technique, have gained more attention recently for speech generation. Codec models represent each audio frame as a sequence of tokens, i.e., discrete embeddings. The discrete and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-31 Alexander H. Liu , Qirui Wang , Yuan Gong , James Glass

We introduce AudioLM, a framework for high-quality audio generation with long-term consistency. AudioLM maps the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this representation…