English
Related papers

Related papers: ManiCM: Real-time 3D Diffusion Policy via Consiste…

200 papers

Diffusion Purification, purifying noised images with diffusion models, has been widely used for enhancing certified robustness via randomized smoothing. However, existing frameworks often grapple with the balance between efficiency and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Yiquan Li , Zhongzhu Chen , Kun Jin , Jiongxiao Wang , Bo Li , Chaowei Xiao

We present a diffusion-based model recipe for real-world control of a highly dexterous humanoid robotic hand, designed for sample-efficient learning and smooth fine-motor action inference. Our system features a newly designed 16-DoF…

Kinematic sensors are often used to analyze movement behaviors in sports and daily activities due to their ease of use and lack of spatial restrictions, unlike video-based motion capturing systems. Still, the generation, and especially the…

Machine Learning · Computer Science 2025-11-27 Heiko Oppel , Michael Munz

We present Diffuse-CLoC, a guided diffusion framework for physics-based look-ahead control that enables intuitive, steerable, and physically realistic motion generation. While existing kinematics motion generation with diffusion models…

The pose-guided person image generation task requires synthesizing photorealistic images of humans in arbitrary poses. The existing approaches use generative adversarial networks that do not necessarily maintain realistic textures or need…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Ankan Kumar Bhunia , Salman Khan , Hisham Cholakkal , Rao Muhammad Anwer , Jorma Laaksonen , Mubarak Shah , Fahad Shahbaz Khan

Prior approaches injecting camera control into diffusion models have focused on specific subsets of 4D consistency tasks: novel view synthesis, text-to-video with camera control, image-to-video, amongst others. Therefore, these fragmented…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Xiang Fan , Sharath Girish , Vivek Ramanujan , Chaoyang Wang , Ashkan Mirzaei , Petr Sushko , Aliaksandr Siarohin , Sergey Tulyakov , Ranjay Krishna

Robot manipulation in the real world is fundamentally constrained by the visual sim2real gap, where depth observations collected in simulation fail to reflect the complex noise patterns inherent to real sensors. In this work, inspired by…

Robotics · Computer Science 2025-12-09 Xiujian Liang , Jiacheng Liu , Mingyang Sun , Qichen He , Cewu Lu , Jianhua Sun

We present FlightDiffusion, a diffusion-model-based framework for training autonomous drones from first-person view (FPV) video. Our model generates realistic video sequences from a single frame, enriched with corresponding action spaces to…

Recently, diffusion models have achieved significant advances in vision, text, and robotics. However, they still face slow generation speeds due to sequential denoising processes. To address this, a parallel sampling method based on Picard…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Junhyuk So , Jiwoong Shin , Chaeyeon Jang , Eunhyeok Park

We propose MonoSE(3)-Diffusion, a monocular SE(3) diffusion framework that formulates markerless, image-based robot pose estimation as a conditional denoising diffusion process. The framework consists of two processes: a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Kangjian Zhu , Haobo Jiang , Yigong Zhang , Jianjun Qian , Jian Yang , Jin Xie

Robotic manipulation in high-precision tasks is essential for numerous industrial and real-world applications where accuracy and speed are required. Yet current diffusion-based policy learning methods generally suffer from low computational…

Robotics · Computer Science 2025-06-23 Sen Wang , Le Wang , Sanping Zhou , Jingyi Tian , Jiayi Li , Haowen Sun , Wei Tang

We propose RoHM, an approach for robust 3D human motion reconstruction from monocular RGB(-D) videos in the presence of noise and occlusions. Most previous approaches either train neural networks to directly regress motion in 3D or learn…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Siwei Zhang , Bharat Lal Bhatnagar , Yuanlu Xu , Alexander Winkler , Petr Kadlecek , Siyu Tang , Federica Bogo

Diffusion policies excel at visuomotor control but often fail catastrophically under severe out-of-distribution (OOD) disturbances, such as unexpected object displacements or visual corruptions. To address this vulnerability, we introduce…

Robotics · Computer Science 2026-03-24 Ziou Hu , Xiangtong Yao , Yuan Meng , Zhenshan Bing , Alois Knoll

Road user trajectory prediction in dynamic environments is a challenging but crucial task for various applications, such as autonomous driving. One of the main challenges in this domain is the multimodal nature of future trajectories…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Younwoo Choi , Ray Coden Mercurius , Soheil Mohamad Alizadeh Shabestary , Amir Rasouli

We introduce Linearly Constrained Diffusion Implicit Models (CDIM), a fast and accurate approach to solving noisy linear inverse problems using diffusion models. Traditional diffusion-based inverse methods rely on numerous projection steps…

Machine Learning · Computer Science 2025-12-01 Vivek Jayaram , Ira Kemelmacher-Shlizerman , Steven M. Seitz , John Thickstun

Diffusion-based generative models have recently emerged as powerful solutions for high-quality synthesis in multiple domains. Leveraging the bidirectional Markov chains, diffusion probabilistic models generate samples by inferring the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Mengyi Zhao , Mengyuan Liu , Bin Ren , Shuling Dai , Nicu Sebe

We propose a novel diffusion-based action model for robotic motion planning. Commonly, established numerical planning approaches are used to solve general motion planning problems, but have significant runtime requirements. By leveraging…

Robotics · Computer Science 2025-09-17 Lennart Clasmeier , Jan-Gerrit Habekost , Connor Gäde , Philipp Allgeuer , Stefan Wermter

Recent large vision-language-action models pretrained on diverse robot datasets have demonstrated the potential for generalizing to new environments with a few in-domain data. However, those approaches usually predict individual discretized…

Robotics · Computer Science 2025-03-25 Zhi Hou , Tianyi Zhang , Yuwen Xiong , Hengjun Pu , Chengyang Zhao , Ronglei Tong , Yu Qiao , Jifeng Dai , Yuntao Chen

Diffusion-based methods have been acknowledged as a powerful paradigm for end-to-end visuomotor control in robotics. Most existing approaches adopt a Diffusion Policy in U-Net architecture (DP-U), which, while effective, suffers from…

Robotics · Computer Science 2025-09-30 Linzhi Wu , Aoran Mei , Xiyue Wang , Guo-Niu Zhu , Zhongxue Gan

Striking a balance between efficiency and transparent motion is a core challenge in human-robot collaboration, as highly expressive movements often incur unnecessary time and energy costs. In collaborative environments, legibility allows a…

Robotics · Computer Science 2026-05-07 Adrien Jacquet Crétides , Mouad Abrini , Hamed Rahimi , Mohamed Chetouani