中文
相关论文

相关论文: D-SCo: Dual-Stream Conditional Diffusion for Monoc…

200 篇论文

Diffusion models have recently gained traction as a powerful class of deep generative priors, excelling in a wide range of image restoration tasks due to their exceptional ability to model data distributions. To solve image restoration…

图像与视频处理 · 电气工程与系统科学 2025-06-10 Xiang Li , Soo Min Kwon , Shijun Liang , Ismail R. Alkhouri , Saiprasad Ravishankar , Qing Qu

A central goal in AI is to represent scenes as compositions of discrete objects, enabling fine-grained, controllable image and video generation. Yet leading diffusion models treat images holistically and rely on text conditioning, creating…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Adil Kaan Akan

Diffusion Denoising models demonstrated impressive results across generative Computer Vision tasks, but they still fail to outperform standard autoregressive solutions in the discrete domain, and only match them at best. In this work, we…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Jia Cheng Hu , Roberto Cavicchioli , Alessandro Capotondi

Although diffusion models can generate high-quality human images, their applications are limited by the instability in generating hands with correct structures. In this paper, we introduce RHanDS, a conditional diffusion-based framework…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Chengrui Wang , Pengfei Liu , Min Zhou , Ming Zeng , Xubin Li , Tiezheng Ge , Bo zheng

The world around us is full of soft objects we perceive and deform with dexterous hand movements. For a robotic hand to control soft objects, it has to acquire online state feedback of the deforming object. While RGB-D cameras can collect…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Elham Amin Mansour , Hehui Zheng , Robert K. Katzschmann

Source localization is the inverse problem of graph information dissemination and has broad practical applications. However, the inherent intricacy and uncertainty in information dissemination pose significant challenges, and the ill-posed…

机器学习 · 计算机科学 2023-04-19 Bosong Huang , Weihao Yu , Ruzhong Xie , Jing Xiao , Jin Huang

We propose RoHM, an approach for robust 3D human motion reconstruction from monocular RGB(-D) videos in the presence of noise and occlusions. Most previous approaches either train neural networks to directly regress motion in 3D or learn…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Siwei Zhang , Bharat Lal Bhatnagar , Yuanlu Xu , Alexander Winkler , Petr Kadlecek , Siyu Tang , Federica Bogo

Anomaly detection, the technique of identifying abnormal samples using only normal samples, has attracted widespread interest in industry. Existing one-model-per-category methods often struggle with limited generalization capabilities due…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Jiawei Zhan , Jinxiang Lai , Bin-Bin Gao , Jun Liu , Xiaochen Chen , Chengjie Wang

Capturing images is a key part of automation for high-level tasks such as scene text recognition. Low-light conditions pose a challenge for high-level perception stacks, which are often optimized on well-lit, artifact-free images.…

图像与视频处理 · 电气工程与系统科学 2023-11-01 Cindy M. Nguyen , Eric R. Chan , Alexander W. Bergman , Gordon Wetzstein

Denoising diffusion models produce high-fidelity image samples by capturing the image distribution in a progressive manner while initializing with a simple distribution and compounding the distribution complexity. Although these models have…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Ayantika Das , Moitreya Chaudhuri , Koushik Bhat , Keerthi Ram , Mihail Bota , Mohanasankar Sivaprakasam

Despite the impressive results of arbitrary image-guided style transfer methods, text-driven image stylization has recently been proposed for transferring a natural image into a stylized one according to textual descriptions of the target…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Nisha Huang , Yuxin Zhang , Fan Tang , Chongyang Ma , Haibin Huang , Yong Zhang , Weiming Dong , Changsheng Xu

Diffusion models have recently advanced video editing, yet controllable editing remains challenging due to the need for precise manipulation of diverse object properties. Current methods require different control signal for diverse editing…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Yuqing Chen , Junjie Wang , Lin Liu , Ruihang Chu , Xiaopeng Zhang , Qi Tian , Yujiu Yang

Multi-turn compositional image generation (M-CIG) is a challenging task that aims to iteratively manipulate a reference image given a modification text. While most of the existing methods for M-CIG are based on generative adversarial…

计算机视觉与模式识别 · 计算机科学 2023-11-15 Chao Wang

Object-centric learning aims to represent visual data with a set of object entities (a.k.a. slots), providing structured representations that enable systematic generalization. Leveraging advanced architectures like Transformers, recent…

计算机视觉与模式识别 · 计算机科学 2023-09-25 Ziyi Wu , Jingyu Hu , Wuyue Lu , Igor Gilitschenski , Animesh Garg

In this paper we consider the problem of acoustic inversion in the context of the optoacoustic tomography image reconstruction problem. By leveraging the ability of the recently proposed diffusion models for image generative tasks among…

图像与视频处理 · 电气工程与系统科学 2024-04-17 M. G. González , M. Vera , A. Dreszman , L. J. Rey Vega

Blind image separation (BIS) refers to the inverse problem of simultaneously estimating and restoring multiple independent source images from a single observation image under conditions of unknown mixing mode and without prior knowledge of…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Jingwei Li , Wei Pu

Denoising diffusion models achieved impressive results on several image generation tasks often outperforming GAN based models. Recently, the generative capabilities of diffusion models have been employed for perceptual image compression,…

图像与视频处理 · 电气工程与系统科学 2025-05-20 Jonas Brenig , Radu Timofte

Diffusion models have demonstrated remarkable efficacy in generating high-quality samples. Existing diffusion-based image restoration algorithms exploit pre-trained diffusion models to leverage data priors, yet they still preserve elements…

图像与视频处理 · 电气工程与系统科学 2024-08-07 Hongjie Wu , Linchao He , Mingqin Zhang , Dongdong Chen , Kunming Luo , Mengting Luo , Ji-Zhe Zhou , Hu Chen , Jiancheng Lv

Disentangling visual layers in real-world images is a persistent challenge in vision and graphics, as such layers often involve non-linear and globally coupled interactions, including shading, reflection, and perspective distortion. In this…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Zheng Gu , Min Lu , Zhida Sun , Dani Lischinski , Daniel Cohen-Or , Hui Huang

Deepfakes pose significant security and privacy threats through malicious facial manipulations. While robust watermarking can aid in authenticity verification and source tracking, existing methods often lack the sufficient robustness…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Chen Sun , Haiyang Sun , Zhiqing Guo , Yunfeng Diao , Liejun Wang , Dan Ma , Gaobo Yang , Keqin Li