English
Related papers

Related papers: Blind Audio Bandwidth Extension: A Diffusion-Based…

200 papers

In this paper the problem of blind super-resolution of sparse signals using arbitrary sampling scheme and atomic lift is discussed. After comprehensive description on blind superresolution problem, it is shown that using Prolate Spheroidal…

Signal Processing · Electrical Eng. & Systems 2019-07-09 Hoomaan Hezaveh , Milad Javadzadeh , MohammadHossein Kahaei

Depth completion, predicting dense depth maps from sparse depth measurements, is an ill-posed problem requiring prior knowledge. Recent methods adopt learning-based approaches to implicitly capture priors, but the priors primarily fit…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Lee Hyoseok , Kyeong Seon Kim , Kwon Byung-Ki , Tae-Hyun Oh

Rabies remains a major public health concern across many African and Asian countries, where accurate diagnosis is critical for effective epidemiological surveillance. The gold standard diagnostic methods rely heavily on fluorescence…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Khalil Akremi , Mariem Handous , Zied Bouslama , Farah Bassalah , Maryem Jebali , Mariem Hanachi , Ines Abdeljaoued-Tej

Pan-sharpening involves reconstructing missing high-frequency information in multi-spectral images with low spatial resolution, using a higher-resolution panchromatic image as guidance. Although the inborn connection with frequency domain,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 Xuanhua He , Keyu Yan , Rui Li , Chengjun Xie , Jie Zhang , Man Zhou

Editing signals using large pre-trained models, in a zero-shot manner, has recently seen rapid advancements in the image domain. However, this wave has yet to reach the audio domain. In this paper, we explore two zero-shot editing…

Sound · Computer Science 2024-05-30 Hila Manor , Tomer Michaeli

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Helin Wang , Jiarui Hai , Yen-Ju Lu , Karan Thakkar , Mounya Elhilali , Najim Dehak

A new method for the design of linear-phase robust far-field broadband beamformers using constrained optimization is proposed. In the method, the maximum passband ripple and minimum stopband attenuation are ensured to be within prescribed…

Systems and Control · Computer Science 2015-06-18 R. C. Nongpiur , D. J. Shpak

We introduce MAVE (Mamba with Cross-Attention for Voice Editing and Synthesis), a novel autoregressive architecture for text-conditioned voice editing and high-fidelity text-to-speech (TTS) synthesis, built on a cross-attentive Mamba…

Sound · Computer Science 2025-10-07 Baher Mohammad , Magauiya Zhussip , Stamatios Lefkimmiatis

Audio zooming, a signal processing technique, enables selective focusing and enhancement of sound signals from a specified region, attenuating others. While traditional beamforming and neural beamforming techniques, centered on creating a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-23 Meng Yu , Dong Yu

The SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-26 Chenglin Xu , Wei Rao , Eng Siong Chng , Haizhou Li

Detecting the presence of animal vocalisations in nature is essential to study animal populations and their behaviors. A recent development in the field is the introduction of the task known as few-shot bioacoustic sound event detection,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-28 Jinhua Liang , Ines Nolasco , Burooj Ghani , Huy Phan , Emmanouil Benetos , Dan Stowell

It's assumed that training data is sufficient in base session of few-shot class-incremental audio classification. However, it's difficult to collect abundant samples for model training in base session in some practical scenarios due to the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Yongjie Si , Yanxiong Li , Jialong Li , Jiaxin Tan , Qianhua He

We describe a joint acoustic echo cancellation (AEC) and blind source extraction (BSE) approach for multi-microphone acoustic frontends. The proposed algorithm blindly estimates AEC and beamforming filters by maximizing the statistical…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-11 Thomas Haubner , Zbyněk Koldovský , Walter Kellermann

Diffusion-based representation learning has achieved substantial attention due to its promising capabilities in latent representation and sample generation. Recent studies have employed an auxiliary encoder to identify a corresponding…

Machine Learning · Computer Science 2025-03-11 Yeongmin Kim , Kwanghyeon Lee , Minsang Park , Byeonghu Na , Il-Chul Moon

In this work we address the problem of blindly reconstructing compressively sensed signals by exploiting the co-sparse analysis model. In the analysis model it is assumed that a signal multiplied by an analysis operator results in a sparse…

Information Theory · Computer Science 2013-03-27 Julian Wörmann , Simon Hawe , Martin Kleinsteuber

Unified image restoration is a significantly challenging task in low-level vision. Existing methods either make tailored designs for specific tasks, limiting their generalizability across various types of degradation, or rely on training…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Huaqiu Li , Yong Wang , Tongwen Huang , Hailang Huang , Haoqian Wang , Xiangxiang Chu

Diffusion models have emerged as powerful tools for solving inverse problems due to their exceptional ability to model complex prior distributions. However, existing methods predominantly assume known forward operators (i.e., non-blind),…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Weimin Bai , Siyi Chen , Wenzheng Chen , He Sun

Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive,…

Recent research has shown that text-to-image diffusion models are capable of generating high-quality images guided by text prompts. But can they be used to generate or approximate real-world images from the seed noise? This is known as the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Weiming Chen , Qifan Liu , Siyi Liu , Yushun Tang , Yijia Wang , Zhihan Zhu , Zhihai He

Directly sending audio signals from a transmitter to a receiver across a noisy channel may absorb consistent bandwidth and be prone to errors when trying to recover the transmitted bits. On the contrary, the recent semantic communication…

Sound · Computer Science 2023-09-15 Eleonora Grassucci , Christian Marinoni , Andrea Rodriguez , Danilo Comminiello