English
Related papers

Related papers: FlexiAct: Towards Flexible Action Control in Heter…

200 papers

To synthesize a realistic action sequence based on a single human image, it is crucial to model both motion patterns and diversity in the action video. This paper proposes an Action Conditional Temporal Variational AutoEncoder (ACT-VAE) to…

Computer Vision and Pattern Recognition · Computer Science 2021-08-13 Xiaogang Xu , Yi Wang , Liwei Wang , Bei Yu , Jiaya Jia

Recent pose-transfer methods aim to generate temporally consistent and fully controllable videos of human action where the motion from a reference video is reenacted by a new identity. We evaluate three state-of-the-art pose-transfer…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Vaclav Knapp , Matyas Bohacek

Facial Appearance Editing (FAE) aims to modify physical attributes, such as pose, expression and lighting, of human facial images while preserving attributes like identity and background, showing great importance in photograph. In spite of…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Qilin Wang , Jiangning Zhang , Chengming Xu , Weijian Cao , Ying Tai , Yue Han , Yanhao Ge , Hong Gu , Chengjie Wang , Yanwei Fu

We present MaskAdapt, a framework for flexible motion adaptation in physics-based humanoid control. The framework follows a two-stage residual learning paradigm. In the first stage, we train a mask-invariant base policy using stochastic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Soomin Park , Eunseong Lee , Kwang Bin Lee , Sung-Hee Lee

A person's movement or relative positioning can be effectively captured by different types of sensors and corresponding sensor output can be utilized in various manipulative techniques for the classification of different human activities.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Utsab Saha , Sawradip Saha , Tahmid Kabir , Shaikh Anowarul Fattah , Mohammad Saquib

Recent self-supervised video representation learning methods focus on maximizing the similarity between multiple augmented views from the same video and largely rely on the quality of generated views. However, most existing methods lack a…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Jinhyung Kim , Taeoh Kim , Minho Shim , Dongyoon Han , Dongyoon Wee , Junmo Kim

We address the task of unsupervised retargeting of human actions from one video to another. We consider the challenging setting where only a few frames of the target is available. The core of our approach is a conditional generative model…

Computer Vision and Pattern Recognition · Computer Science 2020-03-26 Jessica Lee , Deva Ramanan , Rohit Girdhar

We introduce a novel text-to-pose video editing method, ReimaginedAct. While existing video editing tasks are limited to changes in attributes, backgrounds, and styles, our method aims to predict open-ended human action changes in video.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Lan Wang , Vishnu Boddeti , Sernam Lim

Current face reenactment and swapping methods mainly rely on GAN frameworks, but recent focus has shifted to pre-trained diffusion models for their superior generation capabilities. However, training these models is resource-intensive, and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Yue Han , Junwei Zhu , Keke He , Xu Chen , Yanhao Ge , Wei Li , Xiangtai Li , Jiangning Zhang , Chengjie Wang , Yong Liu

Zero-shot skeleton-based action recognition aims to develop models capable of identifying actions beyond the categories encountered during training. Previous approaches have primarily focused on aligning visual and semantic representations…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Wenhan Wu , Zhishuai Guo , Chen Chen , Hongfei Xue , Aidong Lu

Face performance capture and reenactment techniques use multiple cameras and sensors, positioned at a distance from the face or mounted on heavy wearable devices. This limits their applications in mobile and outdoor environments. We present…

Computer Vision and Pattern Recognition · Computer Science 2019-05-28 Mohamed Elgharib , Mallikarjun BR , Ayush Tewari , Hyeongwoo Kim , Wentao Liu , Hans-Peter Seidel , Christian Theobalt

Temporal Action Segmentation (TAS) is an essential task in video analysis, aiming to segment and classify continuous frames into distinct action segments. However, the ambiguous boundaries between actions pose a significant challenge for…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Shuaibing Wang , Shunli Wang , Mingcheng Li , Dingkang Yang , Haopeng Kuang , Ziyun Qian , Lihua Zhang

With the rapid advancement of 2D generative models, preserving subject identity while enabling diverse editing has emerged as a critical research focus. Existing methods typically face inherent trade-offs between identity preservation and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Linyan Huang , Haonan Lin , Yanning Zhou , Kaiwen Xiao

Large language model (LLM) agents typically rely on reactive decision-making paradigms such as ReAct, selecting actions conditioned on growing execution histories. While effective for short tasks, these approaches often lead to redundant…

Artificial Intelligence · Computer Science 2026-03-02 Yihan , Wen , Xin Chen

Robust perception and dynamics modeling are fundamental to real-world robotic policy learning. Recent methods employ video diffusion models (VDMs) to enhance robotic policies, improving their understanding and modeling of the physical…

The essence of a video lies in its dynamic motions, including character actions, object movements, and camera movements. While text-to-video generative diffusion models have recently advanced in creating diverse contents, controlling…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Yuxin Zhang , Fan Tang , Nisha Huang , Haibin Huang , Chongyang Ma , Weiming Dong , Changsheng Xu

Controllable text-to-image (T2I) diffusion models generate images conditioned on both text prompts and semantic inputs of other modalities like edge maps. Nevertheless, current controllable T2I methods commonly face challenges related to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Xuehai He , Jian Zheng , Jacob Zhiyuan Fang , Robinson Piramuthu , Mohit Bansal , Vicente Ordonez , Gunnar A Sigurdsson , Nanyun Peng , Xin Eric Wang

Learning robot manipulation from human videos is appealing due to the scale and diversity of human demonstrations, but transferring such demonstrations to executable robot behavior remains challenging. Prior work either relies on robot data…

Robotics · Computer Science 2026-05-05 Yifan Han , Jianxiang Liu , Haoyu Zhang , Yuqi Gu , Yunhan Guo , Wenzhao Lian

We present MOFA-Video, an advanced controllable image animation method that generates video from the given image using various additional controllable signals (such as human landmarks reference, manual trajectories, and another even…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Muyao Niu , Xiaodong Cun , Xintao Wang , Yong Zhang , Ying Shan , Yinqiang Zheng

Vision-based robot learning often relies on dense image or point-cloud inputs, which are computationally heavy and entangle irrelevant background features. Existing keypoint-based approaches can focus on manipulation-centric features and be…

Robotics · Computer Science 2026-04-17 Anukriti Singh , Kasra Torshizi , Khuzema Habib , Kelin Yu , Ruohan Gao , Pratap Tokekar