English
Related papers

Related papers: Diffusion for De-Occlusion: Accessory-Aware Diffus…

200 papers

Deep learning has the potential to enhance speech signals and increase their intelligibility for users of hearing aids. Deep models suited for real-world application should feature a low computational complexity and low processing delay of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-31 Nils L. Westhausen , Hendrik Kayser , Theresa Jansen , Bernd T. Meyer

Generic image inpainting aims to complete a corrupted image by borrowing surrounding information, which barely generates novel content. By contrast, multi-modal inpainting provides more flexible and useful controls on the inpainted content,…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Shaoan Xie , Zhifei Zhang , Zhe Lin , Tobias Hinz , Kun Zhang

Room Impulse Responses estimation is a fundamental problem in spatial audio processing and speech enhancement. In this paper, we build upon our previously introduced diffusion-based inpainting framework for Room Impulse Response…

Sound · Computer Science 2026-03-31 Sagi Della Torre , Mirco Pezzoli , Fabio Antonacci , Sharon Gannot

Face parsing infers a pixel-wise label map for each semantic facial component. Previous methods generally work well for uncovered faces, however, they overlook facial occlusion and ignore some contextual areas outside a single face,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Jianhua Qiua , Weihua Liu , Chaochao Lin , Jiaojiao Li , Haoping Yu , Said Boumaraf

Inpainting, the process of filling missing or corrupted image parts, has broad applications in medical imaging. However, generating anatomically accurate synthetic polyp images for clinical AI is a largely underexplored problem. In…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Duy-Bao Bui , Hoang-Khang Nguyen , Thao Thi Phuong Dao , Kim Anh Phung , Tam V. Nguyen , Justin Zhan , Minh-Triet Tran , Trung-Nghia Le

With Neural Radiance Fields (NeRFs) arising as a powerful 3D representation, research has investigated its various downstream tasks, including inpainting NeRFs with 2D images. Despite successful efforts addressing the view consistency and…

Image and Video Processing · Electrical Eng. & Systems 2025-04-04 Jingyu Shi , Achleshwar Luthra , Jiazhi Li , Xiang Gao , Xiyun Song , Zongfang Lin , David Gu , Heather Yu

Diffusion models have shown a great ability at bridging the performance gap between predictive and generative approaches for speech enhancement. We have shown that they may even outperform their predictive counterparts for non-additive…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-13 Jean-Marie Lemercier , Julius Richter , Simon Welker , Timo Gerkmann

The audio denoising technique has captured widespread attention in the deep neural network field. Recently, the audio denoising problem has been converted into an image generation task, and deep learning-based approaches have been applied…

Sound · Computer Science 2024-06-14 Junhui Li , Pu Wang , Jialu Li , Youshan Zhang

Robust invisible watermarking aims to embed hidden messages into images such that they survive various manipulations while remaining imperceptible. However, powerful diffusion-based image generation and editing models now enable realistic…

Cryptography and Security · Computer Science 2025-11-11 Wenkai Fu , Finn Carter , Yue Wang , Emily Davis , Bo Zhang

Creating in-silico data with generative AI promises a cost-effective alternative to staining, imaging, and annotating whole slide images in computational pathology. Diffusion models are the state-of-the-art solution for generating in-silico…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Dominik Winter , Nicolas Triltsch , Marco Rosati , Anatoliy Shumilov , Ziya Kokaragac , Yuri Popov , Thomas Padel , Laura Sebastian Monasor , Ross Hill , Markus Schick , Nicolas Brieu

Adjusting transparency is a common method of mitigating occlusion but is often detrimental for understanding the relative depth relationships between objects as well as removes potentially important information from the occluding object. We…

Human-Computer Interaction · Computer Science 2025-07-01 George Bell , Alma Cantu

Although recent speech processing technologies have achieved significant improvements in objective metrics, there still remains a gap in human perceptual quality. This paper proposes Diffiner, a novel solution that utilizes the powerful…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-11 Masato Hirano , Ryosuke Sawata , Naoki Murata , Shusuke Takahashi , Yuki Mitsufuji

Speech-driven 3D facial animation is important for many multimedia applications. Recent work has shown promise in using either Diffusion models or Transformer architectures for this task. However, their mere aggregation does not lead to…

Computer Vision and Pattern Recognition · Computer Science 2024-02-09 Zhiyuan Ma , Xiangyu Zhu , Guojun Qi , Chen Qian , Zhaoxiang Zhang , Zhen Lei

The use of supervised deep learning techniques to detect pathologies in brain MRI scans can be challenging due to the diversity of brain anatomy and the need for annotated data sets. An alternative approach is to use unsupervised anomaly…

Image and Video Processing · Electrical Eng. & Systems 2023-03-08 Finn Behrendt , Debayan Bhattacharya , Julia Krüger , Roland Opfer , Alexander Schlaefer

Since acquiring large amounts of realistic blurry-sharp image pairs is difficult and expensive, learning blind image deblurring from unpaired data is a more practical and promising solution. Unfortunately, dominant approaches rely heavily…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Chengxu Liu , Lu Qi , Jinshan Pan , Xueming Qian , Ming-Hsuan Yang

In this paper, we present PGDI, a diffusion-based speech inpainting framework for restoring missing or severely corrupted speech segments. Unlike previous methods that struggle with speaker variability or long gap lengths, PGDI can…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-13 Mordehay Moradi , Sharon Gannot

Image inpainting is the task of reconstructing missing or damaged parts of an image in a way that seamlessly blends with the surrounding content. With the advent of advanced generative models, especially diffusion models and generative…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Xingzhong Hou , Jie Wu , Boxiao Liu , Yi Zhang , Guanglu Song , Yunpeng Liu , Yu Liu , Haihang You

With the rapid development of mobile devices and the fast increase of sensitive data, secure and convenient mobile authentication technologies are desired. Except for traditional passwords, many mobile devices have biometric-based…

Sound · Computer Science 2025-04-02 Yadong Xie , Fan Li , Yue Wu , Yu Wang

Recognizing the expressions of partially occluded faces is a challenging computer vision problem. Previous expression recognition methods, either overlooked this issue or resolved it using extreme assumptions. Motivated by the fact that the…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Hui Ding , Peng Zhou , Rama Chellappa

Diffusion models have recently emerged as powerful generative models in medical imaging. However, it remains a major challenge to combine these data-driven models with domain knowledge to guide brain imaging problems. In neuroimaging,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Ana Lawry Aguila , Dina Zemlyanker , You Cheng , Sudeshna Das , Daniel C. Alexander , Oula Puonti , Annabel Sorby-Adams , W. Taylor Kimberly , Juan Eugenio Iglesias
‹ Prev 1 4 5 6 7 8 10 Next ›