English
Related papers

Related papers: D-SCo: Dual-Stream Conditional Diffusion for Monoc…

200 papers

We propose residual denoising diffusion models (RDDM), a novel dual diffusion process that decouples the traditional single denoising diffusion process into residual diffusion and noise diffusion. This dual diffusion framework expands the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Jiawei Liu , Qiang Wang , Huijie Fan , Yinong Wang , Yandong Tang , Liangqiong Qu

Computer vision techniques play a central role in the perception stack of autonomous vehicles. Such methods are employed to perceive the vehicle surroundings given sensor data. 3D LiDAR sensors are commonly used to collect sparse 3D point…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Lucas Nunes , Rodrigo Marcuzzi , Benedikt Mersch , Jens Behley , Cyrill Stachniss

The practical applications of diffusion models have been limited by the misalignment between generated images and corresponding text prompts. Recent studies have introduced direct preference optimization (DPO) to enhance the alignment of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Zijing Hu , Fengda Zhang , Kun Kuang

While diffusion models have demonstrated remarkable progress in 2D image generation and editing, extending these capabilities to 3D editing remains challenging, particularly in maintaining multi-view consistency. Classical approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Yufeng Chi , Huimin Ma , Kafeng Wang , Jianmin Li

Recent years have seen significant progress in human image generation, particularly with the advancements in diffusion models. However, existing diffusion methods encounter challenges when producing consistent hand anatomy and the generated…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Anton Pelykh , Ozge Mercanoglu Sincan , Richard Bowden

Decompositional reconstruction of 3D scenes, with complete shapes and detailed texture of all objects within, is intriguing for downstream applications but remains challenging, particularly with sparse views as input. Recent approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Junfeng Ni , Yu Liu , Ruijie Lu , Zirui Zhou , Song-Chun Zhu , Yixin Chen , Siyuan Huang

The limited understanding capacity of the visual encoder in Contrastive Language-Image Pre-training (CLIP) has become a key bottleneck for downstream performance. This capacity includes both Discriminative Ability (D-Ability), which…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Boyu Han , Qianqian Xu , Shilong Bao , Zhiyong Yang , Ruochen Cui , Xilin Zhao , Qingming Huang

Accurate Speed-of-Sound (SoS) reconstruction from acoustic waveforms is a cornerstone of ultrasound computed tomography (USCT), enabling quantitative velocity mapping that reveals subtle anatomical details and pathological variations often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yujia Wu , Shuoqi Chen , Shiru Wang , Yucheng Tang , Petr Bruza , Geoffrey P. Luke

Hand pose estimation from a single image has many applications. However, approaches to full 3D body pose estimation are typically trained on day-to-day activities or actions. As such, detailed hand-to-hand interactions are poorly…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Maksym Ivashechkin , Oscar Mendez , Richard Bowden

Diffusion models are a new class of generative models, and have dramatically promoted image generation with unprecedented quality and diversity. Existing diffusion models mainly try to reconstruct input image from a corrupted one with a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Ling Yang , Jingwei Liu , Shenda Hong , Zhilong Zhang , Zhilin Huang , Zheming Cai , Wentao Zhang , Bin Cui

We propose a novel 3d colored shape reconstruction method from a single RGB image through diffusion model. Diffusion models have shown great development potentials for high-quality 3D shape generation. However, most existing work based on…

Computer Vision and Pattern Recognition · Computer Science 2023-02-14 Bo Li , Xiaolin Wei , Fengwei Chen , Bin Liu

Automating the synthesis of coordinated bimanual piano performances poses significant challenges, particularly in capturing the intricate choreography between the hands while preserving their distinct kinematic signatures. In this paper, we…

Sound · Computer Science 2025-09-05 Zihao Liu , Mingwen Ou , Zunnan Xu , Jiaqi Huang , Haonan Han , Ronghui Li , Xiu Li

Currently, methods for single-image deblurring based on CNNs and transformers have demonstrated promising performance. However, these methods often suffer from perceptual limitations, poor generalization ability, and struggle with heavy or…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Xiaoyang Liu , Yuquan Wang , Zheng Chen , Jiezhang Cao , He Zhang , Yulun Zhang , Xiaokang Yang

Computational tomography (CT) provides high-resolution medical imaging, but it can expose patients to high radiation. X-ray scanners have low radiation exposure, but their resolutions are low. This paper proposes a new conditional diffusion…

Image and Video Processing · Electrical Eng. & Systems 2025-01-20 Yun Su Jeong , Hye Bin Yoo , Il Yong Chun

Although diffusion methods excel in text-to-image generation, generating accurate hand gestures remains a major challenge, resulting in severe artifacts, such as incorrect number of fingers or unnatural gestures. To enable the diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Qifan Fu , Xu Chen , Muhammad Asad , Shanxin Yuan , Changjae Oh , Gregory Slabaugh

The emergence of diffusion models has significantly advanced generative AI, improving the quality, realism, and creativity of image and video generation. Among them, Stable Diffusion (StableDiff) stands out as a key model for text-to-image…

Hardware Architecture · Computer Science 2025-07-03 Zhican Wang , Guanghui He , Hongxiang Fan

Diffusion models, known for their powerful generative capabilities, play a crucial role in addressing real-world super-resolution challenges. However, these models often focus on improving local textures while neglecting the impacts of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Chunyang Bi , Xin Luo , Sheng Shen , Mengxi Zhang , Huanjing Yue , Jingyu Yang

Image restoration is essential for enhancing degraded images across computer vision tasks. However, most existing methods address only a single type of degradation (e.g., blur, noise, or haze) at a time, limiting their real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Debabrata Mandal , Soumitri Chattopadhyay , Guansen Tong , Praneeth Chakravarthula

Existing RGB-D salient object detection (SOD) approaches concentrate on the cross-modal fusion between the RGB stream and the depth stream. They do not deeply explore the effect of the depth map itself. In this work, we design a single…

Computer Vision and Pattern Recognition · Computer Science 2020-07-16 Xiaoqi Zhao , Lihe Zhang , Youwei Pang , Huchuan Lu , Lei Zhang

Diffusion models have recently achieved significant success in various image manipulation tasks, including image super-resolution and perceptual quality enhancement. Pretrained text-to-image models, such as Stable Diffusion, have exhibited…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Sanchar Palit , Subhasis Chaudhuri , Biplab Banerjee