English
Related papers

Related papers: Guiding Diffusion-Based Articulated Object Generat…

200 papers

Reconstructing the 3D shape of an object from a single RGB image is a long-standing and highly challenging problem in computer vision. In this paper, we propose a novel method for single-image 3D reconstruction which generates a sparse…

Computer Vision and Pattern Recognition · Computer Science 2023-02-24 Luke Melas-Kyriazi , Christian Rupprecht , Andrea Vedaldi

Text-guided diffusion models have achieved remarkable success in object inpainting by providing high-level semantic guidance through text prompts. However, they often lack precise pixel-level spatial control, especially in scenarios…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Yongle Zhang , Yimin Liu , Yan Huang , Qiang Wu

Current subject-driven image generation methods encounter significant challenges in person-centric image generation. The reason is that they learn the semantic scene and person generation by fine-tuning a common pre-trained diffusion, which…

Computer Vision and Pattern Recognition · Computer Science 2024-05-06 Yibin Wang , Weizhong Zhang , Jianwei Zheng , Cheng Jin

Denoising diffusion models show remarkable performances in generative tasks, and their potential applications in perception tasks are gaining interest. In this paper, we introduce a novel framework named DiffRef3D which adopts the diffusion…

Computer Vision and Pattern Recognition · Computer Science 2023-10-26 Se-Ho Kim , Inyong Koo , Inyoung Lee , Byeongjun Park , Changick Kim

Generating background scenes for salient objects plays a crucial role across various domains including creative design and e-commerce, as it enhances the presentation and context of subjects by integrating them into tailored environments.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Amir Erfan Eshratifar , Joao V. B. Soares , Kapil Thadani , Shaunak Mishra , Mikhail Kuznetsov , Yueh-Ning Ku , Paloma de Juan

Diffusion MRI (dMRI) is an important neuroimaging technique with high acquisition costs. Deep learning approaches have been used to enhance dMRI and predict diffusion biomarkers through undersampled dMRI. To generate more comprehensive raw…

Image and Video Processing · Electrical Eng. & Systems 2024-07-11 Juanhua Zhang , Ruodan Yan , Alessandro Perelli , Xi Chen , Chao Li

Diffusion models achieve state-of-the-art performance in generating realistic objects and have been successfully applied to images, text, and videos. Recent work has shown that diffusion can also be defined on graphs, including graph…

Machine Learning · Computer Science 2023-02-09 Alex M. Tseng , Nathaniel Diamant , Tommaso Biancalani , Gabriele Scalia

We present a novel 3D shape completion framework that unifies multimodal conditioning, leveraging both 2D images and 3D partial scans through a latent diffusion model. Shapes are represented as Truncated Signed Distance Functions (TSDFs)…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Simon Schaefer , Juan D. Galvis , Xingxing Zuo , Stefan Leutengger

Understanding and manipulating articulated objects, such as doors and drawers, is crucial for robots operating in human environments. We wish to develop a system that can learn to articulate novel objects with no prior interaction, after…

Robotics · Computer Science 2024-05-03 Harry Zhang , Ben Eisner , David Held

Text-guided diffusion models have shown superior performance in image/video generation and editing. While few explorations have been performed in 3D scenarios. In this paper, we discuss three fundamental and interesting problems on this…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Gang Li , Heliang Zheng , Chaoyue Wang , Chang Li , Changwen Zheng , Dacheng Tao

We introduce a framework for joint grounded scene graph - image generation, a challenging task involving high-dimensional, multi-modal structured data. To effectively model this complex joint distribution, we adopt a factorized approach:…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Bicheng Xu , Qi Yan , Renjie Liao , Lele Wang , Leonid Sigal

In music-driven dance motion generation, most existing methods use hand-crafted features and neglect that music foundation models have profoundly impacted cross-modal content generation. To bridge this gap, we propose a diffusion-based…

Sound · Computer Science 2025-02-28 Xinran Liu , Zhenhua Feng , Diptesh Kanojia , Wenwu Wang

Nine-degrees-of-freedom (9-DoF) object pose and size estimation is crucial for enabling augmented reality and robotic manipulation. Category-level methods have received extensive research attention due to their potential for generalization…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Jian Liu , Wei Sun , Hui Yang , Pengchao Deng , Chongpei Liu , Nicu Sebe , Hossein Rahmani , Ajmal Mian

Creating diverse and high-quality 3D assets with an automatic generative model is highly desirable. Despite extensive efforts on 3D generation, most existing works focus on the generation of a single category or a few categories. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Ziang Cao , Fangzhou Hong , Tong Wu , Liang Pan , Ziwei Liu

We introduce AO-Grasp, a grasp proposal method that generates 6 DoF grasps that enable robots to interact with articulated objects, such as opening and closing cabinets and appliances. AO-Grasp consists of two main contributions: the…

Modeling wind-driven object dynamics from video observations is highly challenging due to the invisibility and spatio-temporal variability of wind, as well as the complex deformations of objects. We present DiffWind, a physics-informed…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yuanhang Lei , Boming Zhao , Zesong Yang , Xingxuan Li , Tao Cheng , Haocheng Peng , Ru Zhang , Yang Yang , Siyuan Huang , Yujun Shen , Ruizhen Hu , Hujun Bao , Zhaopeng Cui

Recently normalizing flows (NFs) have demonstrated state-of-the-art performance on modeling 3D point clouds while allowing sampling with arbitrary resolution at inference time. However, these flow-based models still require long training…

Computer Vision and Pattern Recognition · Computer Science 2021-12-01 Janis Postels , Mengya Liu , Riccardo Spezialetti , Luc Van Gool , Federico Tombari

Applying diffusion models to physically-based material estimation and generation has recently gained prominence. In this paper, we propose \ttt, a novel material reconstruction framework for 3D objects, offering the following advantages.…

Graphics · Computer Science 2025-11-25 Xiuchao Wu , Pengfei Zhu , Jiangjing Lyu , Xinguo Liu , Jie Guo , Yanwen Guo , Weiwei Xu , Chengfei Lyu

3D Gaussian Splatting (3DGS) has demonstrated its advantages in achieving fast and high-quality rendering. As point clouds serve as a widely-used and easily accessible form of 3D representation, bridging the gap between point clouds and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Weiqi Zhang , Junsheng Zhou , Haotian Geng , Wenyuan Zhang , Yu-Shen Liu

With the rising industrial attention to 3D virtual modeling technology, generating novel 3D content based on specified conditions (e.g. text) has become a hot issue. In this paper, we propose a new generative 3D modeling framework called…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Muheng Li , Yueqi Duan , Jie Zhou , Jiwen Lu