English
Related papers

Related papers: From Generated Human Videos to Physically Plausibl…

200 papers

Current methods for generating human motion videos rely on extracting pose sequences from reference videos, which restricts flexibility and control. Additionally, due to the limitations of pose detection techniques, the extracted pose…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Yuhang Zhang , Yuan Zhou , Zeyu Liu , Yuxuan Cai , Qiuyue Wang , Aidong Men , Huan Yang

Generative videos have the potential to revolutionize game development by autonomously creating new content. In this paper, we present GameFactory, a framework for action-controlled scene-generalizable game video generation. We first…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Jiwen Yu , Yiran Qin , Xintao Wang , Pengfei Wan , Di Zhang , Xihui Liu

Text-driven human motion generation is an emerging task in animation and humanoid robot design. Existing algorithms directly generate the full sequence which is computationally expensive and prone to errors as it does not pay special…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Zichen Geng , Caren Han , Zeeshan Hayder , Jian Liu , Mubarak Shah , Ajmal Mian

Training general-purpose robots requires learning from large and diverse data sources. Current approaches rely heavily on teleoperated demonstrations which are difficult to scale. We present a scalable framework for training manipulation…

Robotics · Computer Science 2026-05-29 Marion Lepert , Jiaying Fang , Jeannette Bohg

Learning robot control policies from human videos is a promising direction for scaling up robot learning. However, how to extract action knowledge (or action representations) from videos for policy learning remains a key challenge. Existing…

Robotics · Computer Science 2025-06-05 Zhao-Heng Yin , Sherry Yang , Pieter Abbeel

Recent advances in data-driven reinforcement learning and motion tracking have substantially improved humanoid locomotion, yet critical practical challenges remain. In particular, while low-level motion tracking and trajectory-following…

Robotics · Computer Science 2026-02-25 Chenxi Han , Yuheng Min , Zihao Huang , Ao Hong , Hang Liu , Yi Cheng , Houde Liu

Video-based human motion transfer creates video animations of humans following a source motion. Current methods show remarkable results for tightly-clad subjects. However, the lack of temporally consistent handling of plausible clothing…

Computer Vision and Pattern Recognition · Computer Science 2020-12-22 Moritz Kappel , Vladislav Golyanik , Mohamed Elgharib , Jann-Ole Henningson , Hans-Peter Seidel , Susana Castillo , Christian Theobalt , Marcus Magnor

Humanoid Reaction Synthesis is pivotal for creating highly interactive and empathetic robots that can seamlessly integrate into human environments, enhancing the way we live, work, and communicate. However, it is difficult to learn the…

Robotics · Computer Science 2024-04-02 Yunze Liu , Changxi Chen , Chenjing Ding , Li Yi

Recent advances have shown that video generation models can enhance robot learning by deriving effective robot actions through inverse dynamics. However, these methods heavily depend on the quality of generated data and struggle with…

Robotics · Computer Science 2025-08-18 Kelin Yu , Sheng Zhang , Harshit Soora , Furong Huang , Heng Huang , Pratap Tokekar , Ruohan Gao

We aim to tackle the interesting yet challenging problem of generating videos of diverse and natural human motions from prescribed action categories. The key issue lies in the ability to synthesize multiple distinct motion sequences that…

Computer Vision and Pattern Recognition · Computer Science 2021-12-21 Chuan Guo , Xinxin Zuo , Sen Wang , Xinshuang Liu , Shihao Zou , Minglun Gong , Li Cheng

We propose a simple and novel method for generating 3D human motion from complex natural language sentences, which describe different velocity, direction and composition of all kinds of actions. Different from existing methods that use…

Computer Vision and Pattern Recognition · Computer Science 2023-04-17 Zhiyuan Ren , Zhihong Pan , Xin Zhou , Le Kang

Humanoid robots hold great promise in assisting humans in diverse environments and tasks, due to their flexibility and adaptability leveraging human-like morphology. However, research in humanoid robots is often bottlenecked by the costly…

Robotics · Computer Science 2024-06-21 Carmelo Sferrazza , Dun-Ming Huang , Xingyu Lin , Youngwoon Lee , Pieter Abbeel

We aim to enable humanoid robots to efficiently solve new manipulation tasks from a few video examples. In-context learning (ICL) is a promising framework for achieving this goal due to its test-time data efficiency and rapid adaptability.…

Building generalist robots capable of performing functional grasping in everyday, open-world environments remains a significant challenge due to the vast diversity of objects and tasks. Existing methods are either constrained to narrow…

Robotics · Computer Science 2026-04-10 Chao Tang , Jiacheng Xu , Haofei Lu , Bolin Zou , Wenlong Dong , Hong Zhang , Danica Kragic

Visual imitation learning enables robotic agents to acquire skills by observing expert demonstration videos. In the one-shot setting, the agent generates a policy after observing a single expert demonstration without additional fine-tuning.…

Robotics · Computer Science 2026-01-01 Raktim Gautam Goswami , Prashanth Krishnamurthy , Yann LeCun , Farshad Khorrami

A novel concept of vision-based intelligent control of robotic arms is developed here in this work. This work enables the controlling of robotic arms motion only with visual inputs, that is, controlling by showing the videos of correct…

Computer Vision and Pattern Recognition · Computer Science 2021-01-28 Debarati B. Chakraborty , Mukesh Sharma , Bhaskar Vijay

Recent advances in video generation models have sparked interest in world models capable of simulating realistic environments. While navigation has been well-explored, physically meaningful interactions that mimic real-world forces remain…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Nate Gillman , Charles Herrmann , Michael Freeman , Daksh Aggarwal , Evan Luo , Deqing Sun , Chen Sun

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Renda Li , Xiaohua Qi , Qiang Ling , Jun Yu , Ziyi Chen , Peng Chang , Mei HanJing Xiao

In video understanding tasks, particularly those involving human motion, synthetic data generation often suffers from uncanny features, diminishing its effectiveness for training. Tasks such as sign language translation, gesture…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Vaclav Knapp , Matyas Bohacek

Humanoid robots are well suited for human habitats due to their morphological similarity, but developing controllers for them is a challenging task that involves multiple sub-problems, such as control, planning and perception. In this…

Robotics · Computer Science 2023-10-11 K. Niranjan Kumar , Irfan Essa , Sehoon Ha