English
Related papers

Related papers: Moving by Looking: Towards Vision-Driven Avatar Mo…

200 papers

A long-standing goal in computer vision is to capture, model, and realistically synthesize human behavior. Specifically, by learning from data, our goal is to enable virtual humans to navigate within cluttered indoor scenes and naturally…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Mohamed Hassan , Duygu Ceylan , Ruben Villegas , Jun Saito , Jimei Yang , Yi Zhou , Michael Black

Assistive visual navigation systems for visually impaired individuals have become increasingly popular thanks to the rise of mobile computing. Most of these devices work by translating visual information into voice commands. In complex…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Hao Wang , Jiayou Qin , Xiwen Chen , Ashish Bastola , John Suchanek , Zihao Gong , Abolfazl Razi

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans actively coordinate…

Robotics · Computer Science 2026-03-11 Justin Yu , Yide Shentu , Di Wu , Pieter Abbeel , Ken Goldberg , Philipp Wu

Egocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologies, and enabling adaptive healthcare solutions by accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Chaitanya Patel , Hiroki Nakamura , Yuta Kyuragi , Kazuki Kozuka , Juan Carlos Niebles , Ehsan Adeli

Rapid progress in deep reinforcement learning has made it increasingly feasible to train controllers for high-dimensional humanoid bodies. However, methods that use pure reinforcement learning with simple reward functions tend to produce…

Robotics · Computer Science 2017-07-11 Josh Merel , Yuval Tassa , Dhruva TB , Sriram Srinivasan , Jay Lemmon , Ziyu Wang , Greg Wayne , Nicolas Heess

This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to data scarcity. We propose Humanoid-VLA, a novel framework…

We bring together ideas from recent work on feature design for egocentric action recognition under one framework by exploring the use of deep convolutional neural networks (CNN). Recent work has shown that features such as hand appearance,…

Computer Vision and Pattern Recognition · Computer Science 2016-05-13 Minghuang Ma , Haoqi Fan , Kris M. Kitani

The formation of eyes led to the big bang of evolution. The dynamics changed from a primitive organism waiting for the food to come into contact for eating food being sought after by visual sensors. The human eye is one of the most…

Computer Vision and Pattern Recognition · Computer Science 2022-06-14 Varun Ravi Kumar

This paper investigates different approaches to build and use digital human avatars toward interactive Virtual Co-presence (VCP) environments. We evaluate the evolution of technologies for creating VCP environments and how the advancement…

Human-Computer Interaction · Computer Science 2022-01-13 Matthew Korban , Xin Li

Humans do not passively observe the visual world -- we actively look in order to act. Motivated by this principle, we introduce EyeRobot, a robotic system with gaze behavior that emerges from the need to complete real-world tasks. We…

Robotics · Computer Science 2025-09-16 Justin Kerr , Kush Hari , Ethan Weber , Chung Min Kim , Brent Yi , Tyler Bonnen , Ken Goldberg , Angjoo Kanazawa

While progress in 2D generative models of human appearance has been rapid, many applications require 3D avatars that can be animated and rendered. Unfortunately, most existing methods for learning generative models of 3D humans with diverse…

Computer Vision and Pattern Recognition · Computer Science 2023-05-04 Zijian Dong , Xu Chen , Jinlong Yang , Michael J. Black , Otmar Hilliges , Andreas Geiger

Well-designed indoor scenes should prioritize how people can act within a space rather than merely what objects to place. However, existing 3D scene generation methods emphasize visual and semantic plausibility, while insufficiently…

Human-Computer Interaction · Computer Science 2026-03-04 Semin Jin , Donghyuk Kim , Jeongmin Ryu , Kyung Hoon Hyun

The rising demand for creating lifelike avatars in the digital realm has led to an increased need for generating high-quality human videos guided by textual descriptions and poses. We propose Dancing Avatar, designed to fabricate human…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Bosheng Qin , Wentao Ye , Qifan Yu , Siliang Tang , Yueting Zhuang

Learning to solve precision-based manipulation tasks from visual feedback using Reinforcement Learning (RL) could drastically reduce the engineering efforts required by traditional robot systems. However, performing fine-grained motor…

Robotics · Computer Science 2022-01-21 Rishabh Jangir , Nicklas Hansen , Sambaran Ghosal , Mohit Jain , Xiaolong Wang

Autonomously navigating a robot in everyday crowded spaces requires solving complex perception and planning challenges. When using only monocular image sensor data as input, classical two-dimensional planning approaches cannot be used.…

Robotics · Computer Science 2022-03-24 Daniel Dugas , Olov Andersson , Roland Siegwart , Jen Jen Chung

Reconstructing dynamic humans interacting with real-world environments from monocular videos is an important and challenging task. Despite considerable progress in 4D neural rendering, existing approaches either model dynamic scenes…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Wenqing Wang , Haosen Yang , Josef Kittler , Xiatian Zhu

The aim our work is to create virtual humans as intelligent entities, which includes approximate the maximum as possible the virtual agent animation to the natural human behavior. In order to accomplish this task, our agent must be capable…

Multiagent Systems · Computer Science 2010-04-27 F. Cherif , R. Chighoub

Tracking 3D human motion from egocentric multi-camera headset is challenged by severe egomotion, partial visibility or occlusions and lack of training data. Existing methods designed for monocular video often require static or slowly-moving…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Nan Yang , Julian Straub , Fan Zhang , Richard Newcombe , Jakob Engel , Lingni Ma

Learning-based approaches to monocular motion capture have recently shown promising results by learning to regress in a data-driven manner. However, due to the challenges in data collection and network designs, it remains challenging for…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Yuxiang Zhang , Hongwen Zhang , Liangxiao Hu , Jiajun Zhang , Hongwei Yi , Shengping Zhang , Yebin Liu

Recent self-supervised learning (SSL) models trained on human-like egocentric visual inputs substantially underperform on image recognition tasks compared to humans. These models train on raw, uniform visual inputs collected from…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Timothy Schaumlöffel , Arthur Aubret , Gemma Roig , Jochen Triesch