English
Related papers

Related papers: ActCam: Zero-Shot Joint Camera and 3D Motion Contr…

200 papers

Recent advancements in human motion synthesis have focused on specific types of motions, such as human-scene interaction, locomotion or human-human interaction, however, there is a lack of a unified system capable of generating a diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Jianqi Chen , Panwen Hu , Xiaojun Chang , Zhenwei Shi , Michael Kampffmeyer , Xiaodan Liang

Estimating camera motion and intrinsics from casual videos is a core challenge in computer vision. Traditional bundle-adjustment based methods, such as SfM and SLAM, struggle to perform reliably on arbitrary data. Although specialized SfM…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Felix Wimbauer , Weirong Chen , Dominik Muhle , Christian Rupprecht , Daniel Cremers

We present visual action prompts, a unified action representation for action-to-video generation of complex high-DoF interactions while maintaining transferable visual dynamics across domains. Action-driven video generation faces a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Yuang Wang , Chao Wen , Haoyu Guo , Sida Peng , Minghan Qin , Hujun Bao , Xiaowei Zhou , Ruizhen Hu

Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D scenes. Replicating such capability, termed compositional 3D reconstruction, is pivotal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Mingyu Dong , Chong Xia , Mingyuan Jia , Weichen Lyu , Long Xu , Zheng Zhu , Yueqi Duan

We introduce InteractPro, a comprehensive framework for dynamic motion-aware image composition. At its core is InteractPlan, an intelligent planner that leverages a Large Vision Language Model (LVLM) for scenario analysis and object…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Weijing Tao , Xiaofeng Yang , Miaomiao Cui , Guosheng Lin

Image generation technology can synthesize condition-specific images to supplement real-world industrial anomaly data and enhance anomaly detection model performance. Existing generation techniques rarely account for the pose and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Zebei Tong , Hongchang Chen , Yujie Lei , Gang Chen , Yushi Liu , Zhi Zheng , Hao Chen , Jieming Zhang , Ying Li , Dongpu Cao

We introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that generates realistic, context-aware dynamics for both the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Di Chang , Hongyi Xu , You Xie , Yipeng Gao , Zhengfei Kuang , Shengqu Cai , Chenxu Zhang , Guoxian Song , Chao Wang , Yichun Shi , Zeyuan Chen , Shijie Zhou , Linjie Luo , Gordon Wetzstein , Mohammad Soleymani

Existing video deraining methods are often trained on paired datasets, either synthetic, which limits their ability to generalize to real-world rain, or captured by static cameras, which restricts their effectiveness in dynamic scenes with…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Tuomas Varanka , Juan Luis Gonzalez , Hyeongwoo Kim , Pablo Garrido , Xu Yao

Recent video generative models have demonstrated impressive visual fidelity, yet they often struggle with semantic, geometric, and identity consistency. In this paper, we propose a system-level framework, termed the Divide-and-Conquer…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Haoyu Zhao , Yuang Zhang , Junqi Cheng , Jiaxi Gu , Zenghui Lu , Peng Shu , Zuxuan Wu , Yu-Gang Jiang

In the field of 3D content generation, single image scene reconstruction methods still struggle to simultaneously ensure the quality of individual assets and the coherence of the overall scene in complex environments, while texture editing…

Graphics · Computer Science 2026-02-18 Xiang Tang , Ruotong Li , Xiaopeng Fan

Single-view indoor scene generation plays a crucial role in a range of real-world applications. However, generating a complete 360{\deg} scene from a single image remains a highly ill-posed and challenging problem. Recent approaches have…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Runmao Yao , Junsheng Zhou , Zhen Dong , Yu-Shen Liu

Generating controllable character animation from a reference image and motion guidance remains a challenging task due to the inherent difficulty of injecting appearance and motion cues into video diffusion models. Prior works often rely on…

Graphics · Computer Science 2025-07-03 Guian Fang , Yuchao Gu , Mike Zheng Shou

Video generation technologies are developing rapidly and have broad potential applications. Among these technologies, camera control is crucial for generating professional-quality videos that accurately meet user expectations. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Wanquan Feng , Jiawei Liu , Pengqi Tu , Tianhao Qi , Mingzhen Sun , Tianxiang Ma , Songtao Zhao , Siyu Zhou , Qian He

Recent advancements in video diffusion models have shown exceptional abilities in simulating real-world dynamics and maintaining 3D consistency. This progress inspires us to investigate the potential of these models to ensure dynamic…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Jianhong Bai , Menghan Xia , Xintao Wang , Ziyang Yuan , Xiao Fu , Zuozhu Liu , Haoji Hu , Pengfei Wan , Di Zhang

Despite remarkable progress in Vision-Language-Action models (VLAs) for robot manipulation, these large pre-trained models require fine-tuning to be deployed in specific environments. These fine-tuned models are highly sensitive to camera…

Robotics · Computer Science 2026-03-09 Hyeongjun Heo , Seungyeon Woo , Sang Min Kim , Junho Kim , Junho Lee , Yonghyeon Lee , Young Min Kim

Due to the emergence of Generative Adversarial Networks, video synthesis has witnessed exceptional breakthroughs. However, existing methods lack a proper representation to explicitly control the dynamics in videos. Human pose, on the other…

Computer Vision and Pattern Recognition · Computer Science 2018-07-31 Ceyuan Yang , Zhe Wang , Xinge Zhu , Chen Huang , Jianping Shi , Dahua Lin

In this paper, we consider a novel and practical case for talking face video generation. Specifically, we focus on the scenarios involving multi-people interactions, where the talking context, such as audience or surroundings, is present.…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Meidai Xuanyuan , Yuwang Wang , Honglei Guo , Qionghai Dai

Recent text-to-video diffusion models have achieved impressive progress. In practice, users often desire the ability to control object motion and camera movement independently for customized video creation. However, current methods lack the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Shiyuan Yang , Liang Hou , Haibin Huang , Chongyang Ma , Pengfei Wan , Di Zhang , Xiaodong Chen , Jing Liao

This paper investigates the problem of zero-shot action recognition, in the setting where no training videos with seen actions are available. For this challenging scenario, the current leading approach is to transfer knowledge from the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Carlo Bretti , Pascal Mettes

Video diffusion models provide powerful real-world simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectory-conditioned video generation address this gap but often rely on 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Yang Bai , Liudi Yang , George Eskandar , Fengyi Shen , Mohammad Altillawi , Ziyuan Liu , Gitta Kutyniok