English
Related papers

Related papers: UniPhys: Unified Planner and Controller with Diffu…

200 papers

We introduce a method for generating realistic pedestrian trajectories and full-body animations that can be controlled to meet user-defined goals. We draw on recent advances in guided diffusion modeling to achieve test-time controllability…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Davis Rempe , Zhengyi Luo , Xue Bin Peng , Ye Yuan , Kris Kitani , Karsten Kreis , Sanja Fidler , Or Litany

The rapid progress of Deepfake technology has made face swapping highly realistic, raising concerns about the malicious use of fabricated facial content. Existing methods often struggle to generalize to unseen domains due to the diverse…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Ke Sun , Shen Chen , Taiping Yao , Hong Liu , Xiaoshuai Sun , Shouhong Ding , Rongrong Ji

Diffusion-based planners have gained significant recent attention for their robustness and performance in long-horizon tasks. However, most existing planners rely on a fixed, pre-specified horizon during both training and inference. This…

Robotics · Computer Science 2025-09-16 Ruijia Liu , Ancheng Hou , Shaoyuan Li , Xiang Yin

Recent advancements in Language Models (LMs) have demonstrated strong semantic reasoning capabilities, enabling their application in high-level decision-making for autonomous driving (AD). However, LMs operate over discrete token spaces and…

Robotics · Computer Science 2026-04-02 Fan Ding , Xuewen Luo , Fengze Yang , Bo Yu , HwaHui Tew , Ganesh Krishnasamy , Junn Yong Loo

Nature evolves creatures with a high complexity of morphological and behavioral intelligence, meanwhile computational methods lag in approaching that diversity and efficacy. Co-optimization of artificial creatures' morphology and control in…

Diffusion models demonstrate superior performance in capturing complex distributions from large-scale datasets, providing a promising solution for quadrupedal locomotion control. However, the robustness of the diffusion planner is…

Models for natural language and images benefit from data scaling behavior: the more data fed into the model, the better they perform. This 'better with more' phenomenon enables the effectiveness of large-scale pre-training on vast amounts…

Machine Learning · Computer Science 2025-10-08 Wenzhuo Tang , Haitao Mao , Danial Dervovic , Ivan Brugere , Saumitra Mishra , Yuying Xie , Jiliang Tang

Sequence modeling approaches have shown promising results in robot imitation learning. Recently, diffusion models have been adopted for behavioral cloning in a sequence modeling fashion, benefiting from their exceptional capabilities in…

Robotics · Computer Science 2024-01-12 Xiang Li , Varun Belagali , Jinghuan Shang , Michael S. Ryoo

Learning from unstructured and uncurated data has become the dominant paradigm for generative approaches in language and vision. Such unstructured and unguided behavior data, commonly known as play, is also easier to collect in robotics but…

Robotics · Computer Science 2023-12-08 Lili Chen , Shikhar Bahl , Deepak Pathak

Despite recent progress, medical foundation models still struggle to unify visual understanding and generation, as these tasks have inherently conflicting goals: semantic abstraction versus pixel-level reconstruction. Existing approaches,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Ruiheng Zhang , Jingfeng Yao , Huangxuan Zhao , Hao Yan , Xiao He , Lei Chen , Zhou Wei , Yong Luo , Zengmao Wang , Lefei Zhang , Dacheng Tao , Bo Du

Efficiently predicting motion plans directly from vision remains a fundamental challenge in robotics, where planning typically requires explicit goal specification and task-specific design. Recent vision-language-action (VLA) models infer…

Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Delong Liu , Zhaohui Hou , Mingjie Zhan , Shihao Han , Zhicheng Zhao , Fei Su

Current multimodal medical image fusion typically assumes that source images are of high quality and perfectly aligned at the pixel level. Its effectiveness heavily relies on these conditions and often deteriorates when handling misaligned…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Dayong Su , Yafei Zhang , Huafeng Li , Jinxing Li , Yu Liu

Accurate modeling of surface pressure fields around objects is fundamental to aerodynamic analysis and design. While neural networks have shown promise as efficient alternatives to expensive Computational Fluid Dynamics (CFD) simulations,…

Computational Engineering, Finance, and Science · Computer Science 2026-01-23 Junhong Zou , Zhenxu Sun , Yueqing Wang , Wei Qiu , Zhaoxiang Zhang , Xiangyu Zhu , Zhen Lei

Denoising diffusion models have shown great promise in human motion synthesis conditioned on natural language descriptions. However, integrating spatial constraints, such as pre-defined motion trajectories and obstacles, remains a challenge…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Korrawe Karunratanakul , Konpat Preechakul , Supasorn Suwajanakorn , Siyu Tang

In this paper, we propose a diffusion-based face swapping framework for the first time, called DiffFace, composed of training ID conditional DDPM, sampling with facial guidance, and a target-preserving blending. In specific, in the training…

Computer Vision and Pattern Recognition · Computer Science 2022-12-29 Kihong Kim , Yunho Kim , Seokju Cho , Junyoung Seo , Jisu Nam , Kychul Lee , Seungryong Kim , KwangHee Lee

Recently, diffusion transformers have gained wide attention with its excellent performance in text-to-image and text-to-vidoe models, emphasizing the need for transformers as backbone for diffusion models. Transformer-based models have…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Nithin Gopalakrishnan Nair , Jeya Maria Jose Valanarasu , Vishal M. Patel

The fashion domain encompasses a variety of real-world multimodal tasks, including multimodal retrieval and multimodal generation. The rapid advancements in artificial intelligence generated content, particularly in technologies like large…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Xiangyu Zhao , Yuehan Zhang , Wenlong Zhang , Xiao-Ming Wu

Robotic manipulation in unstructured environments requires planners to reason jointly about free-space motion and sustained, frictional contact with the environment. Existing (local) planning and simulation frameworks typically separate…

Robotics · Computer Science 2026-02-05 Bingkun Huang , Xin Ma , Nilanjan Chakraborty , Riddhiman Laha

Flow matching models have emerged as a strong alternative to diffusion models, but existing inversion and editing methods designed for diffusion are often ineffective or inapplicable to them. The straight-line, non-crossing trajectories of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Guanlong Jiao , Biqing Huang , Kuan-Chieh Wang , Renjie Liao