English
Related papers

Related papers: DragAnything: Motion Control for Anything using En…

200 papers

While generative models have excelled at creating static 3D content, the pursuit of systems that understand how objects move and respond to interactions remains a fundamental challenge. Current methods for articulated motion lie at a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Tianshan Zhang , Zeyu Zhang , Hao Tang

Deep learning-based methods for video pedestrian detection and tracking require large volumes of training data to achieve good performance. However, data acquisition in crowded public environments raises data privacy concerns -- we are not…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Matteo Fabbri , Guillem Braso , Gianluca Maugeri , Orcun Cetintas , Riccardo Gasparini , Aljosa Osep , Simone Calderara , Laura Leal-Taixe , Rita Cucchiara

Recent advances in video generative models enable the synthesis of realistic human-object interaction videos across a wide range of scenarios and object categories, including complex dexterous manipulations that are difficult to capture…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Hyeonwoo Kim , Jeonghwan Kim , Kyungwon Cho , Hanbyul Joo

We propose a new method for realistic human motion transfer using a generative adversarial network (GAN), which generates a motion video of a target character imitating actions of a source character, while maintaining high authenticity of…

Graphics · Computer Science 2023-05-09 Yang-Tian Sun , Qian-Cheng Fu , Yue-Ren Jiang , Zitao Liu , Yu-Kun Lai , Hongbo Fu , Lin Gao

We propose a method to interactively control the animation of fluid elements in still images to generate cinemagraphs. Specifically, we focus on the animation of fluid elements like water, smoke, fire, which have the properties of repeating…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Aniruddha Mahapatra , Kuldeep Kulkarni

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Gen Li , Bo Zhao , Jianfei Yang , Laura Sevilla-Lara

Accurate object geometry estimation is essential for many downstream tasks, including robotic manipulation and physical interaction. Although vision is the dominant modality for shape perception, it becomes unreliable under occlusions or…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Langzhe Gu , Hung-Jui Huang , Mohamad Qadri , Michael Kaess , Wenzhen Yuan

We consider the problem of estimating an object's physical properties such as mass, friction, and elasticity directly from video sequences. Such a system identification problem is fundamentally ill-posed due to the loss of information…

Interacting with real-world objects in Mixed Reality (MR) often proves difficult when they are crowded, distant, or partially occluded, hindering straightforward selection and manipulation. We observe that these difficulties stem from…

Human-Computer Interaction · Computer Science 2025-07-25 Xiaoan Liu , Difan Jia , Xianhao Carton Liu , Mar Gonzalez-Franco , Chen Zhu-Tian

3D editing has shown remarkable capability in editing scenes based on various instructions. However, existing methods struggle with achieving intuitive, localized editing, such as selectively making flowers blossom. Drag-style editing has…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Chenghao Gu , Zhenzhe Li , Zhengqi Zhang , Yunpeng Bai , Shuzhao Xie , Zhi Wang

Inspired by the emergent behaviors in large language models that generalized human intelligence, the research community is pursuing similar emergent capabilities within world models, with a emphasis on modeling the physical world. Within…

Artificial Intelligence · Computer Science 2026-05-21 Kunqi Xu , Jitao Li , Jianglong Ye , Tianshu Tang , Isabella Liu , Sifei Liu , Xueyan Zou

By generating plausible and smooth transitions between two image frames, video inbetweening is an essential tool for video editing and long video synthesis. Traditional works lack the capability to generate complex large motions. While…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Maham Tanveer , Yang Zhou , Simon Niklaus , Ali Mahdavi Amiri , Hao Zhang , Krishna Kumar Singh , Nanxuan Zhao

Synthesizing 3D human avatars interacting realistically with a scene is an important problem with applications in AR/VR, video games and robotics. Towards this goal, we address the task of generating a virtual human -- hands and full body…

Robotics · Computer Science 2023-03-30 Purva Tendulkar , Dídac Surís , Carl Vondrick

Recent advances in video generation have shown promise for generating future scenarios, critical for planning and control in autonomous driving and embodied intelligence. However, real-world applications demand more than visually plausible…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Tianshuo Xu , Zhifei Chen , Leyi Wu , Hao Lu , Yuying Chen , Lihui Jiang , Bingbing Liu , Yingcong Chen

Visual navigation using only a single camera and a topological map has recently become an appealing alternative to methods that require additional sensors and 3D maps. This is typically achieved through an "image-relative" approach to…

The essence of a video lies in its dynamic motions, including character actions, object movements, and camera movements. While text-to-video generative diffusion models have recently advanced in creating diverse contents, controlling…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Yuxin Zhang , Fan Tang , Nisha Huang , Haibin Huang , Chongyang Ma , Weiming Dong , Changsheng Xu

Most tracking-by-detection methods employ a local search window around the predicted object location in the current frame assuming the previous location is accurate, the trajectory is smooth, and the computational capacity permits a search…

Computer Vision and Pattern Recognition · Computer Science 2015-12-01 Gao Zhu , Fatih Porikli , Hongdong Li

Human video generation task has gained significant attention with the advancement of deep generative models. Generating realistic videos with human movements is challenging in nature, due to the intricacies of human body topology and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Zhangsihao Yang , Mengyi Shan , Mohammad Farazi , Wenhui Zhu , Yanxi Chen , Xuanzhao Dong , Yalin Wang

In this paper, we present a diffusion model-based framework for animating people from a single image for a given target 3D motion sequence. Our approach has two core components: a) learning priors about invisible parts of the human body and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Boyi Li , Junming Chen , Jathushan Rajasegaran , Yossi Gandelsman , Alexei A. Efros , Jitendra Malik

Action recognition is a fundamental task in video understanding. Existing methods typically extract unified features to process all actions in one video, which makes it challenging to model the interactions between different objects in…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Tianci Wu , Guangming Zhu , Jiang Lu , Siyuan Wang , Ning Wang , Nuoye Xiong , Zhang Liang
‹ Prev 1 3 4 5 6 7 10 Next ›