English
Related papers

Related papers: An Octave-based Multi-Resolution CQT Architecture …

200 papers

The escalating challenges of managing vast sensor-generated data, particularly in audio applications, necessitate innovative solutions. Current systems face significant computational and storage demands, especially in real-time applications…

Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstruction and…

Sound · Computer Science 2025-02-11 Tushar Dhyani , Florian Lux , Michele Mancusi , Giorgio Fabbro , Fritz Hohl , Ngoc Thang Vu

Recent advancements in music generation have garnered significant attention, yet existing approaches face critical limitations. Some current generative models can only synthesize either the vocal track or the accompaniment track. While some…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-04 Ziqian Ning , Huakang Chen , Yuepeng Jiang , Chunbo Hao , Guobin Ma , Shuai Wang , Jixun Yao , Lei Xie

In this study, we propose a framework for chirp-based communications by exploiting discrete Fourier transform-spread orthogonal frequency division multiplexing (DFT-s-OFDM). We show that a well-designed frequency-domain spectral shaping…

Signal Processing · Electrical Eng. & Systems 2020-11-24 Alphan Sahin , Nozhan Hosseini , Hosseinali Jamal , Safi Shams Muhtasimul Hoque , David W. Matolak

Inspired by recent developments in neural speech coding and diffusion-based language modeling, we tackle speech enhancement by modeling the conditional distribution of clean speech codes given noisy speech codes using absorbing discrete…

Sound · Computer Science 2026-02-27 Philippe Gonzalez

While diffusion models are best known for their performance in generative tasks, they have also been successfully applied to many other tasks, including audio source separation. However, current generative approaches to music source…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-24 Yun-Ning , Hung , Richard Vogl , Filip Korzeniowski , Igor Pereira

Deep learning-based super-resolution (SR) methods often perform pixel-wise computations uniformly across entire images, even in homogeneous regions where high-resolution refinement is redundant. We propose the Quadtree Diffusion Model…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Donglin Yang , Paul Vicol , Xiaojuan Qi , Renjie Liao , Xiaofan Zhang

A novel frequency domain training sequence and the corresponding carrier frequency offset (CFO) estimator are proposed for orthogonal frequency division multiplexing (OFDM) systems over frequency-selective fading channels. The proposed…

Information Theory · Computer Science 2017-03-22 Yanxiang Jiang , Xiqi Gao , Xiaohu You

Recent Diffusion Transformers (DiTs) have shown impressive capabilities in generating high-quality single-modality content, including images, videos, and audio. However, it is still under-explored whether the transformer-based diffuser can…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Kai Wang , Shijian Deng , Jing Shi , Dimitrios Hatzinakos , Yapeng Tian

Diffusion-based methods demonstrate significant potential for remote sensing image super-resolution at large scaling factors, particularly in reference-based super-resolution (RefSR) where high-resolution reference images provide critical…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Bin Luo , Runmin Dong , Zhaoyang Luo , Jinxiao Zhang , Jiyao Zhao , Fan Wei , Haohuan Fu

Magnetic Resonance cardiac diffusion tensor imaging (cDTI) and cardiac intravoxel incoherent motion imaging enables probing of in vivo myofiber architecture and myocardial perfusion surrogates. To study the impact of experimental parameters…

Sparse-View CT (SVCT) reconstruction enhances temporal resolution and reduces radiation dose, yet its clinical use is hindered by artifacts due to view reduction and domain shifts from scanner, protocol, or anatomical variations, leading to…

Image and Video Processing · Electrical Eng. & Systems 2026-04-24 Haodong Li , Shuo Han , Haiyang Mao , Yu Shi , Changsheng Fang , Jianjia Zhang , Weiwen Wu , Hengyong Yu

Diffusion models have demonstrated impressive image synthesis performance, yet many UNet-based models are trained at certain fixed resolutions. Their quality tends to degrade when generating images at out-of-training resolutions. We trace…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Jiaxuan Ren , Junhan Zhu , Huan Wang

We introduce ImmerseDiffusion, an end-to-end generative audio model that produces 3D immersive soundscapes conditioned on the spatial, temporal, and environmental conditions of sound objects. ImmerseDiffusion is trained to generate…

Sound · Computer Science 2025-02-11 Mojtaba Heydari , Mehrez Souden , Bruno Conejo , Joshua Atkins

The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusion models and audio diffusion models, along with a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Bingqi Ma , Linlong Lang , Ming Zhang , Dailan He , Xingtong Ge , Yi Zhang , Guanglu Song , Yu Liu

Recent advances in talking face generation have significantly improved facial animation synthesis. However, existing approaches face fundamental limitations: 3DMM-based methods maintain temporal consistency but lack fine-grained regional…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Kangwei Liu , Junwu Liu , Yun Cao , Jinlin Guo , Xiaowei Yi

Diffusion-based generative models have reformed generative AI, and also enabled new capabilities in the science domain, e.g., fast generation of 3D structures of molecules. In such tasks, there is often a symmetry in the system, identifying…

Machine Learning · Computer Science 2026-05-15 Yixian Xu , Yusong Wang , Shengjie Luo , Kaiyuan Gao , Tianyu He , Di He , Chang Liu

Neural audio compression has emerged as a promising technology for efficiently representing speech, music, and general audio. However, existing methods suffer from significant performance degradation at limited bitrates, where the available…

Sound · Computer Science 2026-05-08 Jin Wang , Wenbin Jiang , Xiangbo Wang , Yubo You , Sheng Fang

Applying diffusion models to physically-based material estimation and generation has recently gained prominence. In this paper, we propose \ttt, a novel material reconstruction framework for 3D objects, offering the following advantages.…

Graphics · Computer Science 2025-11-25 Xiuchao Wu , Pengfei Zhu , Jiangjing Lyu , Xinguo Liu , Jie Guo , Yanwen Guo , Weiwei Xu , Chengfei Lyu

Quantum Error Mitigation is essential for enhancing the reliability of quantum computing experiments. The adaptive KIK error mitigation method has demonstrated significant advantages, including resilience to temporal noise drifts,…

Quantum Physics · Physics 2026-03-10 Ben Bar , Jader P. Santos , Raam Uzdin