中文
相关论文

相关论文: PhysHMR: Learning Humanoid Control Policies from V…

200 篇论文

Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Lixin Xue , Chengwei Zheng , Georgios Paschalidis , Chen Guo , Manuel Kaufmann , Juan Zarate , Dimitrios Tzionas

Human mesh recovery can be approached using either regression-based or optimization-based methods. Regression models achieve high pose accuracy but struggle with model-to-image alignment due to the lack of explicit 2D-3D correspondences. In…

计算机视觉与模式识别 · 计算机科学 2025-02-07 Chongyang Xu , Buzhen Huang , Chengfang Zhang , Ziliang Feng , Yangang Wang

We propose CrossHuman, a novel method that learns cross-guidance from parametric human model and multi-frame RGB images to achieve high-quality 3D human reconstruction. To recover geometry details and texture even in invisible regions, we…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Liliang Chen , Jiaqi Li , Han Huang , Yandong Guo

We present a novel method for populating 3D indoor scenes with virtual humans that can navigate in the environment and interact with objects in a realistic manner. Existing approaches rely on training sequences that contain captured human…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Kaifeng Zhao , Yan Zhang , Shaofei Wang , Thabo Beeler , Siyu Tang

Transferring human motion and appearance between videos of human actors remains one of the key challenges in Computer Vision. Despite the advances from recent image-to-image translation approaches, there are several transferring contexts…

计算机视觉与模式识别 · 计算机科学 2021-04-29 Thiago L. Gomes , Renato Martins , João Ferreira , Rafael Azevedo , Guilherme Torres , Erickson R. Nascimento

Avatars are important to create interactive and immersive experiences in virtual worlds. One challenge in animating these characters to mimic a user's motion is that commercial AR/VR products consist only of a headset and controllers,…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Daniele Reda , Jungdam Won , Yuting Ye , Michiel van de Panne , Alexander Winkler

Regression-based methods have recently shown promising results in reconstructing human meshes from monocular images. By directly mapping raw pixels to model parameters, these methods can produce parametric models in a feed-forward manner…

计算机视觉与模式识别 · 计算机科学 2021-08-25 Hongwen Zhang , Yating Tian , Xinchi Zhou , Wanli Ouyang , Yebin Liu , Limin Wang , Zhenan Sun

Human behavior is fundamentally shaped by visual perception -- our ability to interact with the world depends on actively gathering relevant information and adapting our movements accordingly. Behaviors like searching for objects, reaching,…

机器人学 · 计算机科学 2025-05-20 Zhengyi Luo , Chen Tessler , Toru Lin , Ye Yuan , Tairan He , Wenli Xiao , Yunrong Guo , Gal Chechik , Kris Kitani , Linxi Fan , Yuke Zhu

We present PHORHUM, a novel, end-to-end trainable, deep neural network methodology for photorealistic 3D human reconstruction given just a monocular RGB image. Our pixel-aligned method estimates detailed 3D geometry and, for the first time,…

计算机视觉与模式识别 · 计算机科学 2022-04-20 Thiemo Alldieck , Mihai Zanfir , Cristian Sminchisescu

There are several challenges in developing a model for multi-tasking humanoid control. Reinforcement learning and imitation learning approaches are quite popular in this domain. However, there is a trade-off between the two. Reinforcement…

机器人学 · 计算机科学 2024-06-18 Siddharth Padmanabhan , Kazuki Miyazawa , Takato Horii , Takayuki Nagai

Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms. Existing cross-embodiment approaches try to map human motion to robot actions, but often…

Due to the visual ambiguity, purely kinematic formulations on monocular human motion capture are often physically incorrect, biomechanically implausible, and can not reconstruct accurate interactions. In this work, we focus on exploiting…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Buzhen Huang , Liang Pan , Yuan Yang , Jingyi Ju , Yangang Wang

In this paper, we introduce a method to automatically reconstruct the 3D motion of a person interacting with an object from a single RGB video. Our method estimates the 3D poses of the person and the object, contact positions, and forces…

计算机视觉与模式识别 · 计算机科学 2019-06-18 Zongmian Li , Jiri Sedlar , Justin Carpentier , Ivan Laptev , Nicolas Mansard , Josef Sivic

In endoscopic surgery, surgeons continuously locate the endoscopic view relative to the anatomy by interpreting the evolving visual appearance of the intraoperative scene in the context of their prior knowledge. Vision-based navigation…

We propose RoHM, an approach for robust 3D human motion reconstruction from monocular RGB(-D) videos in the presence of noise and occlusions. Most previous approaches either train neural networks to directly regress motion in 3D or learn…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Siwei Zhang , Bharat Lal Bhatnagar , Yuanlu Xu , Alexander Winkler , Petr Kadlecek , Siyu Tang , Federica Bogo

We present a novel paradigm of building an animatable 3D human representation from a monocular video input, such that it can be rendered in any unseen poses and views. Our method is based on a dynamic Neural Radiance Field (NeRF) rigged by…

计算机视觉与模式识别 · 计算机科学 2022-08-19 Gusi Te , Xiu Li , Xiao Li , Jinglu Wang , Wei Hu , Yan Lu

Learning 3D human motion from 2D inputs is a fundamental task in the realms of computer vision and computer graphics. Many previous methods grapple with this inherently ambiguous task by introducing motion priors into the learning process.…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Shuaiying Hou , Hongyu Tao , Junheng Fang , Changqing Zou , Hujun Bao , Weiwei Xu

We present HSImul3R, a unified framework for simulation-ready 3D reconstruction of human-scene interactions (HSI) from casual captures, including sparse-view images and monocular videos. Existing methods suffer from a perception-simulation…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yukang Cao , Haozhe Xie , Fangzhou Hong , Long Zhuo , Zhaoxi Chen , Liang Pan , Ziwei Liu

High-fidelity digital human representations are increasingly in demand in the digital world, particularly for interactive telepresence, AR/VR, 3D graphics, and the rapidly evolving metaverse. Even though they work well in small spaces,…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Zexu Huang , Sarah Monazam Erfani , Siying Lu , Mingming Gong

Recent advances in 3D foundation models have led to growing interest in reconstructing humans and their surrounding environments. However, most existing approaches focus on monocular inputs, and extending them to multi-view settings…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Sangmin Kim , Minhyuk Hwang , Geonho Cha , Dongyoon Wee , Jaesik Park