English
Related papers

Related papers: MonoMobility: Zero-Shot 3D Mobility Analysis from …

200 papers

Existing deep learning-based approaches for monocular 3D object detection in autonomous driving often model the object as a rotated 3D cuboid while the object's geometric shape has been ignored. In this work, we propose an approach for…

Computer Vision and Pattern Recognition · Computer Science 2021-08-26 Zongdai Liu , Dingfu Zhou , Feixiang Lu , Jin Fang , Liangjun Zhang

Dynamic scene reconstruction is a long-term challenge in the field of 3D vision. Recently, the emergence of 3D Gaussian Splatting has provided new insights into this problem. Although subsequent efforts rapidly extend static 3D Gaussian to…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Ruijie Zhu , Yanzhe Liang , Hanzhi Chang , Jiacheng Deng , Jiahao Lu , Wenfei Yang , Tianzhu Zhang , Yongdong Zhang

Existing methods for reconstructing objects and humans from a monocular image suffer from severe mesh collisions and performance limitations for interacting occluding objects. This paper introduces a method to obtain a globally consistent…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Sarthak Batra , Partha P. Chakrabarti , Simon Hadfield , Armin Mustafa

For the task of mobility analysis of 3D shapes, we propose joint analysis for simultaneous motion part segmentation and motion attribute estimation, taking a single 3D model as input. The problem is significantly different from those…

Computer Vision and Pattern Recognition · Computer Science 2019-03-13 Xiaogang Wang , Bin Zhou , Yahao Shi , Xiaowu Chen , Qinping Zhao , Kai Xu

We present a method for jointly training the estimation of depth, ego-motion, and a dense 3D translation field of objects relative to the scene, with monocular photometric consistency being the sole source of supervision. We show that this…

Computer Vision and Pattern Recognition · Computer Science 2020-11-10 Hanhan Li , Ariel Gordon , Hang Zhao , Vincent Casser , Anelia Angelova

We propose a semantics-driven unsupervised learning approach for monocular depth and ego-motion estimation from videos in this paper. Recent unsupervised learning methods employ photometric errors between synthetic view and actual image as…

Computer Vision and Pattern Recognition · Computer Science 2020-06-09 Xiaobin Wei , Jianjiang Feng , Jie Zhou

We present an on-line 3D visual object tracking framework for monocular cameras by incorporating spatial knowledge and uncertainty from semantic mapping along with high frequency measurements from visual odometry. Using a combination of…

Computer Vision and Pattern Recognition · Computer Science 2016-03-15 Prateek Singhal , Ruffin White , Henrik Christensen

Monocular 3D object detection plays a crucial role in autonomous driving. However, existing monocular 3D detection algorithms depend on 3D labels derived from LiDAR measurements, which are costly to acquire for new datasets and challenging…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Fulong Ma , Xiaoyang Yan , Guoyang Zhao , Xiaojie Xu , Yuxuan Liu , Jun Ma , Ming Liu

The estimation of optical flow and 6-DoF ego-motion, two fundamental tasks in 3D vision, has typically been addressed independently. For neuromorphic vision (e.g., event cameras), however, the lack of robust data association makes solving…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Wenpu Li , Bangyan Liao , Yi Zhou , Qi Xu , Pian Wan , Peidong Liu

We present MOSAIC-GS, a novel, fully explicit, and computationally efficient approach for high-fidelity dynamic scene reconstruction from monocular videos using Gaussian Splatting. Monocular reconstruction is inherently ill-posed due to the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Svitlana Morkva , Maximum Wilder-Smith , Michael Oechsle , Alessio Tonioni , Marco Hutter , Vaishakh Patil

Building high-fidelity digital twins of articulated objects from visual data remains a central challenge. Existing approaches depend on multi-view captures of the object in discrete, static states, which severely constrains their real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Lijun Guo , Haoyu Zhao , Xingyue Zhao , Rong Fu , Linghao Zhuang , Siteng Huang , Zhongyu Li , Hua Zou

3D human motion capture from monocular RGB images respecting interactions of a subject with complex and possibly deformable environments is a very challenging, ill-posed and under-explored problem. Existing methods address it only weakly…

Computer Vision and Pattern Recognition · Computer Science 2022-08-18 Zhi Li , Soshi Shimada , Bernt Schiele , Christian Theobalt , Vladislav Golyanik

This paper proposes GraviCap, i.e., a new approach for joint markerless 3D human motion capture and object trajectory estimation from monocular RGB videos. We focus on scenes with objects partially observed during a free flight. In contrast…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Rishabh Dabral , Soshi Shimada , Arjun Jain , Christian Theobalt , Vladislav Golyanik

This dissertation is a multifaceted contribution to the advancement of vision-based 3D perception technologies. In the first segment, the thesis introduces structural enhancements to both monocular and stereo 3D object detection algorithms.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Yuxuan Liu

Videos of robots interacting with objects encode rich information about the objects' dynamics. However, existing video prediction approaches typically do not explicitly account for the 3D information from videos, such as robot actions and…

Robotics · Computer Science 2024-10-25 Mingtong Zhang , Kaifeng Zhang , Yunzhu Li

City administrations increasingly rely on comprehensive databases and urban digital twins of city assets, such as traffic signs and trees, as well as incidents like graffiti or road damage, to maintain an effective overview of urban…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Miriam Louise Carnot , Jonas Kunze , Erik Quinten Fastermann , Eric Peukert , André Ludwig , Bogdan Franczyk

We present a simple lightweight markerless facial performance capture framework using just a monocular video input that combines Active Appearance Models for feature tracking and prior constraints on 3D shapes into an integrated objective…

Computer Vision and Pattern Recognition · Computer Science 2019-01-17 Shridhar Ravikumar

Template-based 3D object tracking still lacks a high-precision benchmark of real scenes due to the difficulty of annotating the accurate 3D poses of real moving video objects without using markers. In this paper, we present a multi-view…

Computer Vision and Pattern Recognition · Computer Science 2022-03-28 Jiachen Li , Bin Wang , Shiqiang Zhu , Xin Cao , Fan Zhong , Wenxuan Chen , Te Li , Jason Gu , Xueying Qin

Open-vocabulary 3D instance segmentation transcends traditional closed-vocabulary methods by enabling the identification of both previously seen and unseen objects in real-world scenarios. It leverages a dual-modality approach, utilizing…

Computer Vision and Pattern Recognition · Computer Science 2024-08-19 Tri Ton , Ji Woo Hong , SooHwan Eom , Jun Yeop Shim , Junyeong Kim , Chang D. Yoo

In this paper, we propose MoDGS, a new pipeline to render novel views of dy namic scenes from a casually captured monocular video. Previous monocular dynamic NeRF or Gaussian Splatting methods strongly rely on the rapid move ment of input…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Qingming Liu , Yuan Liu , Jiepeng Wang , Xianqiang Lyv , Peng Wang , Wenping Wang , Junhui Hou