中文
相关论文

相关论文: GeoMan: Temporally Consistent Human Geometry Estim…

200 篇论文

This paper explores the capabilities of convolutional neural networks to deal with a task that is easily manageable for humans: perceiving 3D pose of a human body from varying angles. However, in our approach, we are restricted to using a…

计算机视觉与模式识别 · 计算机科学 2017-07-21 Agne Grinciunaite , Amogh Gudi , Emrah Tasli , Marten den Uyl

The goal of video-based person re-identification is to match two input videos, so that the distance of the two videos is small if two videos contain the same person. A common approach for person re-identification is to first extract image…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Tanzila Rahman , Mrigank Rochan , Yang Wang

Reward-based fine-tuning of video diffusion models is an effective approach to improve the quality of generated videos, as it can fine-tune models without requiring real-world video datasets. However, it can sometimes be limited to specific…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Takehiro Aoshima , Yusuke Shinohara , Byeongseon Park

We introduce GeCo, a geometry-grounded metric for jointly detecting geometric deformation and occlusion-inconsistency artifacts in static scenes. By fusing residual motion and depth priors, GeCo produces interpretable, dense consistency…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Leslie Gu , Junhwa Hur , Charles Herrmann , Fangneng Zhan , Todd Zickler , Deqing Sun , Hanspeter Pfister

Accurate height estimation from monocular aerial imagery presents a significant challenge due to its inherently ill-posed nature. This limitation is rooted in the absence of adequate geometric constraints available to the model when…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Xiaomou Hou , Wanshui Gan , Naoto Yokoya

Accurate surround-view depth estimation provides a competitive alternative to laser-based sensors and is essential for 3D scene understanding in autonomous driving. While empirical studies have proposed various approaches that primarily…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Weimin Liu , Wenjun Wang , Joshua H. Meng

Existing deep models predict 2D and 3D kinematic poses from video that are approximately accurate, but contain visible errors that violate physical constraints, such as feet penetrating the ground and bodies leaning at extreme angles. In…

计算机视觉与模式识别 · 计算机科学 2020-07-27 Davis Rempe , Leonidas J. Guibas , Aaron Hertzmann , Bryan Russell , Ruben Villegas , Jimei Yang

Monocular dynamic reconstruction is a challenging and long-standing vision problem due to the highly ill-posed nature of the task. Existing approaches depend on templates, are effective only in quasi-static scenes, or fail to model 3D…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Qianqian Wang , Vickie Ye , Hang Gao , Weijia Zeng , Jake Austin , Zhengqi Li , Angjoo Kanazawa

Reconstructing dynamic 3D scenes (i.e., 4D geometry) from monocular video is an important yet challenging problem. Conventional multiview geometry-based approaches often struggle with dynamic motion, whereas recent learning-based methods…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Jinjie Mai , Wenxuan Zhu , Haozhe Liu , Bing Li , Cheng Zheng , Jürgen Schmidhuber , Bernard Ghanem

We propose SelfRecon, a clothed human body reconstruction method that combines implicit and explicit representations to recover space-time coherent geometries from a monocular self-rotating human video. Explicit methods require a predefined…

计算机视觉与模式识别 · 计算机科学 2022-04-06 Boyi Jiang , Yang Hong , Hujun Bao , Juyong Zhang

We present an approach to generate a 360-degree view of a person with a consistent, high-resolution appearance from a single input image. NeRF and its variants typically require videos or images from different viewpoints. Most existing…

计算机视觉与模式识别 · 计算机科学 2023-11-16 Badour AlBahar , Shunsuke Saito , Hung-Yu Tseng , Changil Kim , Johannes Kopf , Jia-Bin Huang

Human pose estimation is one of the key problems in computer vision that has been studied in the recent years. The significance of human pose estimation is in the higher level tasks of understanding human actions applications such as…

计算机视觉与模式识别 · 计算机科学 2014-08-26 Oinam Binarani Devi , Nissi S. Paul , Y. Jayanta Singh

We present, GauHuman, a 3D human model with Gaussian Splatting for both fast training (1 ~ 2 minutes) and real-time rendering (up to 189 FPS), compared with existing NeRF-based implicit representation modelling frameworks demanding hours of…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Shoukang Hu , Ziwei Liu

The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensics, social media, and exploration. Existing classification-based approaches operate at a…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Parth Parag Kulkarni , Rohit Gupta , Prakash Chandra Chhipa , Mubarak Shah

Reference-driven image completion, which restores missing regions in a target view using additional images, is particularly challenging when the target view differs significantly from the references. Existing generative methods rely solely…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Beibei Lin , Tingting Chen , Robby T. Tan

Creating realistic avatars from a single RGB image is an attractive yet challenging problem. Due to its ill-posed nature, recent works leverage powerful prior from 2D diffusion models pretrained on large datasets. Although 2D diffusion…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Yuxuan Xue , Xianghui Xie , Riccardo Marin , Gerard Pons-Moll

Recent video depth estimation methods achieve great performance by following the paradigm of image depth estimation, i.e., typically fine-tuning pre-trained video diffusion models with massive data. However, we argue that video depth…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Haodong Li , Chen Wang , Jiahui Lei , Kostas Daniilidis , Lingjie Liu

Understanding dynamic scenes from casual videos is critical for scalable robot learning, yet four-dimensional (4D) reconstruction under strictly monocular settings remains highly ill-posed. To address this challenge, our key insight is that…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Can Li , Jie Gu , Jingmin Chen , Fangzhou Qiu , Lei Sun

Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D structure. In monocular settings, however, such constraints are absent, leading to severe…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Daniel Rho , Jun Myeong Choi , Matthew Thornton , Biswadip Dey , Roni Sengupta

Video-conditioned 4D shape generation aims to recover time-varying 3D geometry and view-consistent appearance directly from an input video. In this work, we introduce a native video-to-4D shape generation framework that synthesizes a single…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Jiraphon Yenphraphai , Ashkan Mirzaei , Jianqi Chen , Jiaxu Zou , Sergey Tulyakov , Raymond A. Yeh , Peter Wonka , Chaoyang Wang