English
Related papers

Related papers: Unlocking Temporal Flexibility: Neural Speech Code…

200 papers

Noise robustness remains a critical challenge for deploying neural speech codecs in real-world acoustic scenarios where background noise is often inevitable. A key observation we make is that even slight input noise perturbations can cause…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-14 Rui-Chen Zheng , Yang Ai , Hui-Peng Du , Li-Rong Dai

Neural audio compression has emerged as a promising technology for efficiently representing speech, music, and general audio. However, existing methods suffer from significant performance degradation at limited bitrates, where the available…

Sound · Computer Science 2026-05-08 Jin Wang , Wenbin Jiang , Xiangbo Wang , Yubo You , Sheng Fang

Building on recent advances in video generation, generative video compression has emerged as a new paradigm for achieving visually pleasing reconstructions. However, existing methods exhibit limited exploitation of temporal correlations,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Xiaoyue Ling , Chuqin Zhou , Chunyi Li , Yunuo Chen , Yuan Tian , Guo Lu , Wenjun Zhang

Neural video compression (NVC) technologies have advanced rapidly in recent years, yielding state-of-the-art schemes such as DCVC-RT that offer superior compression efficiency to H.266/VVC and real-time encoding/decoding capabilities.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Hui Xiang , Yifan Bian , Li Li , Jingran Wu , Xianguo Zhang , Dong Liu

Neural Radiance Fields (NeRF) have exhibited highly effective performance for photorealistic novel view synthesis recently. However, the key limitation it meets is the reliance on a hand-crafted frequency annealing strategy to recover 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Rui Qian , Chenyangguang Zhang , Yan Di , Guangyao Zhai , Ruida Zhang , Jiayu Guo , Benjamin Busam , Jian Pu

The wide deployment of speech-based biometric systems usually demands high-performance speaker recognition algorithms. However, most of the prior works for speaker recognition either process the speech in the frequency domain or time…

Sound · Computer Science 2023-03-08 Jiguo Li , Tianzi Zhang , Xiaobin Liu , Lirong Zheng

CTC-based ASR systems face computational and memory bottlenecks in resource-limited environments. Traditional CTC decoders, requiring up to 90% of processing time in systems (e.g., wav2vec2-large on L4 GPUs), face inefficiencies due to…

Machine Learning · Computer Science 2025-10-13 Atul Shree , Harshith Jupuru

Inspired by the facts that retinal cells actually segregate the visual scene into different attributes (e.g., spatial details, temporal motion) for respective neuronal processing, we propose to first decompose the input video into…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Ming Lu , Tong Chen , Dandan Ding , Fengqing Zhu , Zhan Ma

This paper introduces FlowMAC, a novel neural audio codec for high-quality general audio compression at low bit rates based on conditional flow matching (CFM). FlowMAC jointly learns a mel spectrogram encoder, quantizer and decoder. At…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-08 Nicola Pia , Martin Strauss , Markus Multrus , Bernd Edler

The neural radiance fields (NeRF) have advanced the development of 3D volumetric video technology, but the large data volumes they involve pose significant challenges for storage and transmission. To address these problems, the existing…

Multimedia · Computer Science 2024-11-11 Zhiyu Zhang , Guo Lu , Huanxiong Liang , Zhengxue Cheng , Anni Tang , Li Song

Extracting a target source from underdetermined mixtures is challenging for beamforming approaches. Recently proposed time-frequency-bin-wise switching (TFS) and linear combination (TFLC) strategies mitigate this by combining multiple…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Changda Chen , Yichen Yang , Wei Liu , Shoji Makino

Although deep neural networks have facilitated significant progress of neural vocoders in recent years, they usually suffer from intrinsic challenges like opaque modeling, inflexible retraining under different input configurations, and…

Sound · Computer Science 2026-03-11 Andong Li , Tong Lei , Zhihang Sun , Rilin Chen , Xiaodong Li , Dong Yu , Chengshi Zheng

Neuroimaging-based prediction methods for intelligence and cognitive abilities have seen a rapid development in literature. Among different neuroimaging modalities, prediction based on functional connectivity (FC) has shown great promise.…

Neurons and Cognition · Quantitative Biology 2023-07-20 Yang Li , Xin Ma , Raj Sunderraman , Shihao Ji , Suprateek Kundu

This paper presents a novel algorithm for building an automatic speech recognition (ASR) model with imperfect training data. Imperfectly transcribed speech is a prevalent issue in human-annotated speech corpora, which degrades the…

Computation and Language · Computer Science 2023-06-05 Dongji Gao , Matthew Wiesner , Hainan Xu , Leibny Paola Garcia , Daniel Povey , Sanjeev Khudanpur

Most speech enhancement algorithms make use of the short-time Fourier transform (STFT), which is a simple and flexible time-frequency decomposition that estimates the short-time spectrum of a signal. However, the duration of short STFT…

Sound · Computer Science 2015-09-03 Scott Wisdom , Thomas Powers , Les Atlas , James Pitton

Despite a short history, neural image codecs have been shown to surpass classical image codecs in terms of rate-distortion performance. However, most of them suffer from significantly longer decoding times, which hinders the practical…

Image and Video Processing · Electrical Eng. & Systems 2023-05-16 Yixin Gao , Runsen Feng , Zongyu Guo , Zhibo Chen

In this work we propose a novel approach to utilize convolutional neural networks for time series forecasting. The time direction of the sequential data with spatial dimensions $D=1,2$ is considered democratically as the input of a…

Machine Learning · Computer Science 2020-01-13 Matthias Weissenbacher

Modern video codecs have been extensively optimized to preserve perceptual quality, leveraging models of the human visual system. However, in split inference systems-where intermediate features from neural network are transmitted instead of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Md Eimran Hossain Eimon , Ashan Perera , Juan Merlos , Velibor Adzic , Hari Kalva

The sequence length along the time axis is often the dominant factor of the computation in speech processing. Works have been proposed to reduce the sequence length for lowering the computational cost in self-supervised speech models.…

Computation and Language · Computer Science 2023-05-10 Hsuan-Jui Chen , Yen Meng , Hung-yi Lee

Learned image compression has exhibited promising compression performance, but variable bitrates over a wide range remain a challenge. State-of-the-art variable rate methods compromise the loss of model performance and require numerous…

Image and Video Processing · Electrical Eng. & Systems 2023-03-13 Kedeng Tong , Yaojun Wu , Yue Li , Kai Zhang , Li Zhang , Xin Jin
‹ Prev 1 3 4 5 6 7 10 Next ›