English
Related papers

Related papers: Blind Audio Bandwidth Extension: A Diffusion-Based…

200 papers

The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly implemented biometric identification models and can be…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-16 Haibin Wu , Heng-Cheng Kuo , Naijun Zheng , Kuo-Hsuan Hung , Hung-Yi Lee , Yu Tsao , Hsin-Min Wang , Helen Meng

We propose a novel pitch estimation technique called DeepF0, which leverages the available annotated data to directly learns from the raw audio in a data-driven manner. F0 estimation is important in various speech processing and music…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-15 Satwinder Singh , Ruili Wang , Yuanhang Qiu

Achieving robust generalization against unseen attacks remains a challenge in Audio Deepfake Detection (ADD), driven by the rapid evolution of generative models. To address this, we propose a framework centered on hard sample…

Sound · Computer Science 2026-04-30 Bo Cheng , Songjun Cao , Xiaoming Zhang , Jie Chen , Long Ma , Fei Chen

We propose a novel unsupervised singing voice detection method which use single-channel Blind Audio Source Separation (BASS) algorithm as a preliminary step. To reach this goal, we investigate three promising BASS approaches which operate…

Sound · Computer Science 2018-05-04 Dominique Fourer , Geoffroy Peeters

Real-world audio is often degraded by numerous factors. This work presents an audio restoration model tailored for high-res music at 44.1kHz. Our model, Audio-to-Audio Schr\"odinger Bridges (A2SB), is capable of both bandwidth extension…

Unsupervised Anomaly Detection (UAD) techniques aim to identify and localize anomalies without relying on annotations, only leveraging a model trained on a dataset known to be free of anomalies. Diffusion models learn to modify inputs $x$…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Sergio Naval Marimont , Matthew Baugh , Vasilis Siomos , Christos Tzelepis , Bernhard Kainz , Giacomo Tarroni

Diffusion-based text-to-image models have demonstrated remarkable capabilities in generating realistic images, but they raise societal and ethical concerns, such as the creation of unsafe content. While concept editing is proposed to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Ruipeng Wang , Junfeng Fang , Jiaqi Li , Hao Chen , Jie Shi , Kun Wang , Xiang Wang

Binaural speech enhancement (BSE) aims to jointly improve the speech quality and intelligibility of noisy signals received by hearing devices and preserve the spatial cues of the target for natural listening. Existing methods often suffer…

Sound · Computer Science 2025-01-09 Jingyuan Wang , Jie Zhang , Shihao Chen , Miao Sun

Diffusion models are a new class of generative models that have shown outstanding performance in image generation literature. As a consequence, studies have attempted to apply diffusion models to other tasks, such as speech enhancement. A…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-10 Philippe Gonzalez , Zheng-Hua Tan , Jan Østergaard , Jesper Jensen , Tommy Sonne Alstrøm , Tobias May

Empirical Bayes (EB) improves the accuracy of simultaneous inference "by learning from the experience of others" (Efron, 2012). Classical EB theory focuses on latent variables that are iid draws from a fitted prior (Efron, 2019). Modern…

Methodology · Statistics 2025-12-24 Bohan Wu , Eli N. Weinstein , David M. Blei

Zero-shot inference is a powerful paradigm that enables the use of large pretrained models for downstream classification tasks without further training. However, these models are vulnerable to inherited biases that can impact their…

Machine Learning · Computer Science 2024-02-13 Dyah Adila , Changho Shin , Linrong Cai , Frederic Sala

Wake word (WW) spotting is challenging in far-field due to the complexities and variations in acoustic conditions and the environmental interference in signal transmission. A suite of carefully designed and optimized audio front-end (AFE)…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-15 Yixin Gao , Noah D. Stein , Chieh-Chi Kao , Yunliang Cai , Ming Sun , Tao Zhang , Shiv Vitaladevuni

A key challenge in maximizing the benefits of Magnetic Resonance Imaging (MRI) in clinical settings is to accelerate acquisition times without significantly degrading image quality. This objective requires a balance between under-sampling…

Machine Learning · Computer Science 2025-06-23 Jacopo Iollo , Geoffroy Oudoumanessah , Carole Lartizien , Michel Dojat , Florence Forbes

Data augmentation is conventionally used to inject robustness in Speaker Verification systems. Several recently organized challenges focus on handling novel acoustic environments. Deep learning based speech enhancement is a modern solution…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-29 Saurabh Kataria , Phani Sankar Nidadavolu , Jesús Villalba , Najim Dehak

Most recent advances in audio dereverberation focus almost exclusively on speech, leaving percussive and drum signals largely unexplored despite their importance in music production. Percussive dereverberation poses distinct challenges due…

Sound · Computer Science 2026-05-12 Dimos Makris , András Barják , Maximos Kaliakatsos-Papakostas

Diffusion/score-based models have recently emerged as powerful generative priors for solving inverse problems, including accelerated MRI reconstruction. While their flexibility allows decoupling the measurement model from the learned prior,…

Image and Video Processing · Electrical Eng. & Systems 2025-09-15 Yaşar Utku Alçalar , Junno Yun , Mehmet Akçakaya

Blind face restoration is a highly ill-posed problem due to the lack of necessary context. Although existing methods produce high-quality outputs, they often fail to faithfully preserve the individual's identity. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Siyu Liu , Zheng-Peng Duan , Jia OuYang , Jiayi Fu , Hyunhee Park , Zikun Liu , Chun-Le Guo , Chongyi Li

In this paper, we address the challenge of image resolution variation for the Segment Anything Model (SAM). SAM, known for its zero-shot generalizability, exhibits a performance degradation when faced with datasets with varying image sizes.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Yiran Song , Qianyu Zhou , Xiangtai Li , Deng-Ping Fan , Xuequan Lu , Lizhuang Ma

Detecting unknown deepfake manipulations remains one of the most challenging problems in face forgery detection. Current state-of-the-art approaches fail to generalize to unseen manipulations, as they primarily rely on supervised training…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Kaede Shiohara , Toshihiko Yamasaki , Vladislav Golyanik

Acousto-optic sensing provides an alternative approach to traditional microphone arrays by shedding light on the interaction of light with an acoustic field. Sound field reconstruction is a fascinating and advanced technique used in…

Sound · Computer Science 2025-07-01 Phuc Duc Nguyen , Kenji Ishikawa , Noboru Harada , Takehiro Moriya
‹ Prev 1 8 9 10 Next ›