中文
相关论文

相关论文: NeMo: 3D Neural Motion Fields from Multiple Video …

200 篇论文

Accurate and reliable multi-object tracking (MOT) in 3D space is essential for advancing robotics and computer vision applications. However, it remains a significant challenge in monocular setups due to the difficulty of mining 3D…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Zhuohao Yan , Shaoquan Feng , Xingxing Li , Yuxuan Zhou , Chunxi Xia , Shengyu Li

High-fidelity reconstruction of head avatars from monocular videos is highly desirable for virtual human applications, but it remains a challenge in the fields of computer graphics and computer vision. In this paper, we propose a two-phase…

图形学 · 计算机科学 2025-03-31 Pilseo Park , Ze Zhang , Michel Sarkis , Ning Bi , Xiaoming Liu , Yiying Tong

While previous years have seen great progress in the 3D reconstruction of humans from monocular videos, few of the state-of-the-art methods are able to handle loose garments that exhibit large non-rigid surface deformations during…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Chen Guo , Tianjian Jiang , Manuel Kaufmann , Chengwei Zheng , Julien Valentin , Jie Song , Otmar Hilliges

Recovering world-coordinate human motion from monocular videos with humanoid robot retargeting is significant for embodied intelligence and robotics. To avoid complex SLAM pipelines or heavy temporal models, we propose a lightweight,…

机器人学 · 计算机科学 2025-12-29 Zhangzheng Tu , Kailun Su , Shaolong Zhu , Yukun Zheng

Novel view synthesis from monocular videos of dynamic scenes with unknown camera poses remains a fundamental challenge in computer vision and graphics. While recent advances in 3D representations such as Neural Radiance Fields (NeRF) and 3D…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Mengqi Guo , Bo Xu , Yanyan Li , Gim Hee Lee

We present HuMoCon, a novel motion-video understanding framework designed for advanced human behavior analysis. The core of our method is a human motion concept discovery framework that efficiently trains multi-modal encoders to extract…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Qihang Fang , Chengcheng Tang , Bugra Tekin , Shugao Ma , Yanchao Yang

Accurate hand and finger tracking from video has significant clinical applications for monitoring activities of daily living and measuring range of motion, yet monocular video approaches for obtaining hand biomechanics remain…

计算机视觉与模式识别 · 计算机科学 2026-05-12 R. James Cotton , Pouyan Firouzabadi , Wendy Murray

Monocular 3D object detection (Mono3D) is a fundamental computer vision task that estimates an object's class, 3D position, dimensions, and orientation from a single image. Its applications, including autonomous driving, augmented reality,…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Abhinav Kumar

Video-based human motion transfer creates video animations of humans following a source motion. Current methods show remarkable results for tightly-clad subjects. However, the lack of temporally consistent handling of plausible clothing…

There has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. While these tasks rely heavily on known camera poses, the problem of finding such poses using structure-from-motion…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Lily Goli , Sara Sabour , Mark Matthews , Marcus Brubaker , Dmitry Lagun , Alec Jacobson , David J. Fleet , Saurabh Saxena , Andrea Tagliasacchi

Learning-based approaches to monocular motion capture have recently shown promising results by learning to regress in a data-driven manner. However, due to the challenges in data collection and network designs, it remains challenging for…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Yuxiang Zhang , Hongwen Zhang , Liangxiao Hu , Jiajun Zhang , Hongwei Yi , Shengping Zhang , Yebin Liu

We investigate the use of prior knowledge of human and animal movement to learn reusable locomotion skills for real legged robots. Our approach builds upon previous work on imitating human or dog Motion Capture (MoCap) data to learn a…

Monocular egocentric 3D human motion capture remains a significant challenge, particularly under conditions of low lighting and fast movements, which are common in head-mounted device applications. Existing methods that rely on RGB cameras…

计算机视觉与模式识别 · 计算机科学 2025-02-13 Christen Millerdurai , Hiroyasu Akada , Jian Wang , Diogo Luvizon , Alain Pagani , Didier Stricker , Christian Theobalt , Vladislav Golyanik

Prior work for articulated 3D shape reconstruction often relies on specialized sensors (e.g., synchronized multi-camera systems), or pre-built 3D deformable models (e.g., SMAL or SMPL). Such methods are not able to scale to diverse sets of…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Gengshan Yang , Minh Vo , Natalia Neverova , Deva Ramanan , Andrea Vedaldi , Hanbyul Joo

Advances in machine learning and wearable sensors offer new opportunities for capturing and analyzing human movement outside specialized laboratories. Accurate assessment of human movement under real-world conditions is essential for…

Although humans have the innate ability to imagine multiple possible actions from videos, it remains an extraordinary challenge for computers due to the intricate camera movements and montages. Most existing motion generation methods…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Liangdong Qiu , Chengxing Yu , Yanran Li , Zhao Wang , Haibin Huang , Chongyang Ma , Di Zhang , Pengfei Wan , Xiaoguang Han

In this paper we present a novel method to estimate 3D human pose and shape from monocular videos. This task requires directly recovering pixel-alignment 3D human pose and body shape from monocular images or videos, which is challenging due…

计算机视觉与模式识别 · 计算机科学 2023-03-02 Sen Yang , Wen Heng , Gang Liu , Guozhong Luo , Wankou Yang , Gang Yu

We present XFormer, a novel human mesh and motion capture method that achieves real-time performance on consumer CPUs given only monocular images as input. The proposed network architecture contains two branches: a keypoint branch that…

计算机视觉与模式识别 · 计算机科学 2023-05-19 Lihui Qian , Xintong Han , Faqiang Wang , Hongyu Liu , Haoye Dong , Zhiwen Li , Huawei Wei , Zhe Lin , Cheng-Bin Jin

We introduce SAM 3D Body (3DB), a promptable model for single-image full-body 3D human mesh recovery (HMR) that demonstrates state-of-the-art performance, with strong generalization and consistent accuracy in diverse in-the-wild conditions.…

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Zhixue Fang , Xu He , Songlin Tang , Haoxian Zhang , Qingfeng Li , Xiaoqiang Liu , Pengfei Wan , Kun Gai