English
Related papers

Related papers: DisPose: Disentangling Pose Guidance for Controlla…

200 papers

Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. However, current…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Ruiyan Wang , Teng Hu , Kaihui Huang , Zihan Su , Ran Yi , Lizhuang Ma

Recent advancements in human video synthesis have enabled the generation of high-quality videos through the application of stable diffusion models. However, existing methods predominantly concentrate on animating solely the human element…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Jinlin Liu , Kai Yu , Mengyang Feng , Xiefan Guo , Miaomiao Cui

Recovering dense human poses from images plays a critical role in establishing an image-to-surface correspondence between RGB images and the 3D surface of the human body, serving the foundation of rich real-world applications, such as…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Haonan Yan , Jiaqi Chen , Xujie Zhang , Shengkai Zhang , Nianhong Jiao , Xiaodan Liang , Tianxiang Zheng

We present a generative model for controllable person image synthesis,as shown in Figure , which can be applied to pose-guided person image synthesis, $i.e.$, converting the pose of a source person image to the target pose while preserving…

Computer Vision and Pattern Recognition · Computer Science 2020-12-24 Shilong Shen

One-shot video-driven talking face generation aims at producing a synthetic talking video by transferring the facial motion from a video to an arbitrary portrait image. Head pose and facial expression are always entangled in facial motion…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Youxin Pang , Yong Zhang , Weize Quan , Yanbo Fan , Xiaodong Cun , Ying Shan , Dong-ming Yan

Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose tracking have been studied, they rely heavily on strong external priors, such as depth data…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Jisu Shin , Junoh Lee , JunGyu Lee , Inhwan Bae , Dohyeon Lee , Hokyun Im , Youngwoon Lee , Hae-Gon Jeon

In multi-view 3D human pose estimation, models typically rely on images captured simultaneously from different camera views to predict a pose at a specific moment. While providing accurate spatial information, this traditional approach…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Ling Li , Changjie Chen , Yuyan Wang , Jiaqing Lyu , Kenglun Chang , Yiyun Chen , Zhidong Deng

Synthesizing realistic human-object interactions (HOI) in video is challenging due to the complex, instance-specific interaction dynamics of both humans and objects. Incorporating controllability in video generation further adds to the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Wanyue Zhang , Lin Geng Foo , Thabo Beeler , Rishabh Dabral , Christian Theobalt

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Jingyun Liang , Jingkai Zhou , Shikai Li , Chenjie Cao , Lei Sun , Yichen Qian , Weihua Chen , Fan Wang

Deep generative modelling for human body analysis is an emerging problem with many interesting applications. However, the latent space learned by such approaches is typically not interpretable, resulting in less flexibility. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2020-02-18 Rodrigo de Bem , Arnab Ghosh , Thalaiyasingam Ajanthan , Ondrej Miksik , Adnane Boukhayma , N. Siddharth , Philip Torr

The 3D Human Pose Estimation (3D HPE) task uses 2D images or videos to predict human joint coordinates in 3D space. Despite recent advancements in deep learning-based methods, they mostly ignore the capability of coupling accessible texts…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Jinglin Xu , Yijie Guo , Yuxin Peng

In this paper, we are interested in the bottom-up paradigm of estimating human poses from an image. We study the dense keypoint regression framework that is previously inferior to the keypoint detection and grouping framework. Our…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Zigang Geng , Ke Sun , Bin Xiao , Zhaoxiang Zhang , Jingdong Wang

In the era of deep learning, human pose estimation from multiple cameras with unknown calibration has received little attention to date. We show how to train a neural model to perform this task with high precision and minimal latency…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Ben Usman , Andrea Tagliasacchi , Kate Saenko , Avneesh Sud

Creating believable motions for various characters has long been a goal in computer graphics. Current learning-based motion synthesis methods depend on extensive motion datasets, which are often challenging, if not impossible, to obtain. On…

Computer Vision and Pattern Recognition · Computer Science 2023-11-01 Qingqing Zhao , Peizhuo Li , Wang Yifan , Olga Sorkine-Hornung , Gordon Wetzstein

Accurate estimation of 3D human motion from monocular video requires modeling both kinematics (body motion without physical forces) and dynamics (motion with physical forces). To demonstrate this, we present SimPoE, a Simulation-based…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Ye Yuan , Shih-En Wei , Tomas Simon , Kris Kitani , Jason Saragih

Upsampling videos of human activity is an interesting yet challenging task with many potential applications ranging from gaming to entertainment and sports broadcasting. The main difficulty in synthesizing video frames in this setting stems…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Hsuan-I Ho , Xu Chen , Jie Song , Otmar Hilliges

Controllable human image generation (HIG) has numerous real-life applications. State-of-the-art solutions, such as ControlNet and T2I-Adapter, introduce an additional learnable branch on top of the frozen pre-trained stable diffusion (SD)…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Xuan Ju , Ailing Zeng , Chenchen Zhao , Jianan Wang , Lei Zhang , Qiang Xu

In recent years, generative artificial intelligence has achieved significant advancements in the field of image generation, spawning a variety of applications. However, video generation still faces considerable challenges in various…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Yuang Zhang , Jiaxi Gu , Li-Wen Wang , Han Wang , Junqi Cheng , Yuefeng Zhu , Fangyuan Zou

Reconstructing 3D human shape and pose from monocular images is challenging despite the promising results achieved by the most recent learning-based methods. The commonly occurred misalignment comes from the facts that the mapping from…

Computer Vision and Pattern Recognition · Computer Science 2020-12-08 Hongwen Zhang , Jie Cao , Guo Lu , Wanli Ouyang , Zhenan Sun

Human image animation involves generating videos from a character photo, allowing user control and unlocking the potential for video and movie production. While recent approaches yield impressive results using high-quality training data,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Zhenzhi Wang , Yixuan Li , Yanhong Zeng , Youqing Fang , Yuwei Guo , Wenran Liu , Jing Tan , Kai Chen , Tianfan Xue , Bo Dai , Dahua Lin