English
Related papers

Related papers: ExoActor: Exocentric Video Generation as Generaliz…

200 papers

We describe an end-to-end generative approach for the segmentation and recognition of human activities. In this approach, a visual representation based on reduced Fisher Vectors is combined with a structured temporal model for recognition.…

Computer Vision and Pattern Recognition · Computer Science 2016-03-18 Hilde Kuehne , Juergen Gall , Thomas Serre

Understanding the world in first-person view is fundamental in Augmented Reality (AR). This immersive perspective brings dramatic visual changes and unique challenges compared to third-person views. Synthetic data has empowered…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Gen Li , Kaifeng Zhao , Siwei Zhang , Xiaozhong Lyu , Mihai Dusmanu , Yan Zhang , Marc Pollefeys , Siyu Tang

Robotic manipulation is essential for the widespread adoption of robots in industrial and home settings and has long been a focus within the robotics community. Advances in artificial intelligence have introduced promising learning-based…

Robotics · Computer Science 2025-03-04 Kelin Li , Shubham M Wagh , Nitish Sharma , Saksham Bhadani , Wei Chen , Chang Liu , Petar Kormushev

We introduce FlexAvatar, a method for creating high-quality and complete 3D head avatars from a single image. A core challenge lies in the limited availability of multi-view data and the tendency of monocular training to yield incomplete 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Tobias Kirschstein , Simon Giebenhain , Matthias Nießner

Human video generation task has gained significant attention with the advancement of deep generative models. Generating realistic videos with human movements is challenging in nature, due to the intricacies of human body topology and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Zhangsihao Yang , Mengyi Shan , Mohammad Farazi , Wenhui Zhu , Yanxi Chen , Xuanzhao Dong , Yalin Wang

We address the challenging problem of robotic grasping and manipulation in the presence of uncertainty. This uncertainty is due to noisy sensing, inaccurate models and hard-to-predict environment dynamics. We quantify the importance of…

Robotic systems can enhance the amount and repeatability of physically guided motor training. Yet their real-world adoption is limited, partly due to non-intuitive trainer/therapist-trainee/patient interactions. To address this gap, we…

Real-world videos naturally portray complex interactions among distinct physical objects, effectively forming dynamic compositions of visual elements. However, most current video generation models synthesize scenes holistically and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Guofeng Zhang , Angtian Wang , Jacob Zhiyuan Fang , Liming Jiang , Haotian Yang , Alan Yuille , Chongyang Ma

Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of exploiting existing…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Jilan Xu , Yifei Huang , Junlin Hou , Guo Chen , Yuejie Zhang , Rui Feng , Weidi Xie

World models empower model-based agents to interactively explore, reason, and plan within imagined environments for real-world decision-making. However, the high demand for interactivity poses challenges in harnessing recent advancements in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Jialong Wu , Shaofeng Yin , Ningya Feng , Xu He , Dong Li , Jianye Hao , Mingsheng Long

Enabling humanoid robots to clean rooms has long been a pursued dream within humanoid research communities. However, many tasks require multi-humanoid collaboration, such as carrying large and heavy furniture together. Given the scarcity of…

Robotics · Computer Science 2024-10-31 Jiawei Gao , Ziqin Wang , Zeqi Xiao , Jingbo Wang , Tai Wang , Jinkun Cao , Xiaolin Hu , Si Liu , Jifeng Dai , Jiangmiao Pang

This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accomplish a given task. Although recent Vision-Language Models (VLMs) have achieved impressive…

Robotics · Computer Science 2026-04-10 Peiran Xu , Jiaqi Zheng , Yadong Mu

Our comprehension of video streams depicting human activities is naturally multifaceted: in just a few moments, we can grasp what is happening, identify the relevance and interactions of objects in the scene, and forecast what will happen…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Simone Alberto Peirone , Francesca Pistilli , Antonio Alliegro , Tatiana Tommasi , Giuseppe Averta

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Qi Wu , Yubo Zhao , Yifan Wang , Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

Human video generation is becoming an increasingly important task with broad applications in graphics, entertainment, and embodied AI. Despite the rapid progress of video diffusion models (VDMs), their use for general-purpose human video…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Hyelin Nam , Hyojun Go , Byeongjun Park , Byung-Hoon Kim , Hyungjin Chung

We consider the problem of synthesizing multi-action human motion sequences of arbitrary lengths. Existing approaches have mastered motion sequence generation in single action scenarios, but fail to generalize to multi-action and…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Rania Briq , Chuhang Zou , Leonid Pishchulin , Chris Broaddus , Juergen Gall

The evolution of autonomous systems in the context of human-robot interaction systems necessitates a synergy between the continuous perception of the environment and the potential actions to navigate or interact within it. We present…

Robotics · Computer Science 2025-09-18 Timothée Dhaussy , Bassam Jabaian , Fabrice Lefèvre

Text-conditioned motion synthesis has made remarkable progress with the emergence of diffusion models. However, the majority of these motion diffusion models are primarily designed for a single character and overlook multi-human…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Zhenzhi Wang , Jingbo Wang , Yixuan Li , Dahua Lin , Bo Dai

We study instruction-guided editing of egocentric videos for interactive AR applications. While recent AI video editors perform well on third-person footage, egocentric views present unique challenges - including rapid egomotion and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Runjia Li , Moayed Haji-Ali , Ashkan Mirzaei , Chaoyang Wang , Arpit Sahni , Ivan Skorokhodov , Aliaksandr Siarohin , Tomas Jakab , Junlin Han , Sergey Tulyakov , Philip Torr , Willi Menapace

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Yue Gao , Yuan Zhou , Jinglu Wang , Xiao Li , Xiang Ming , Yan Lu
‹ Prev 1 8 9 10 Next ›