中文
相关论文

相关论文: MonoArt: Progressive Structural Reasoning for Mono…

200 篇论文

Given a single image of a general object such as a chair, could we also restore its articulated 3D shape similar to human modeling, so as to animate its plausible articulations and diverse motions? This is an interesting new question that…

计算机视觉与模式识别 · 计算机科学 2022-07-07 Ji Yang , Xinxin Zuo , Sen Wang , Zhenbo Yu , Xingyu Li , Bingbing Ni , Minglun Gong , Li Cheng

Visual SLAM systems targeting static scenes have been developed with satisfactory accuracy and robustness. Dynamic 3D object tracking has then become a significant capability in visual SLAM with the requirement of understanding dynamic…

计算机视觉与模式识别 · 计算机科学 2022-10-06 Hanwei Zhang , Hideaki Uchiyama , Shintaro Ono , Hiroshi Kawasaki

Reconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are built upon multi-view geometry. State-of-the-art (SOTA)…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Wei Yin , Chi Zhang , Hao Chen , Zhipeng Cai , Gang Yu , Kaixuan Wang , Xiaozhi Chen , Chunhua Shen

State-of-the-art learning-based monocular 3D reconstruction methods learn priors over object categories on the training set, and as a result struggle to achieve reasonable generalization to object categories unseen during training. In this…

计算机视觉与模式识别 · 计算机科学 2020-06-30 Miguel Angel Bautista , Walter Talbott , Shuangfei Zhai , Nitish Srivastava , Joshua M Susskind

3D object reconstruction is important for semantic scene understanding. It is challenging to reconstruct detailed 3D shapes from monocular images directly due to a lack of depth information, occlusion and noise. Most current methods…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Ziwei Liao , Steven L. Waslander

Monocular 3D foundation models offer an extensible solution for perception tasks, making them attractive for broader 3D vision applications. In this paper, we propose MoRe, a training-free Monocular Geometry Refinement method designed to…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Dongki Jung , Jaehoon Choi , Yonghan Lee , Sungmin Eum , Heesung Kwon , Dinesh Manocha

Estimating precise metric depth and scene reconstruction from monocular endoscopy is a fundamental task for surgical navigation in robotic surgery. However, traditional stereo matching adopts binocular images to perceive the depth…

机器人学 · 计算机科学 2022-11-29 Ruofeng Wei , Bin Li , Hangjie Mo , Fangxun Zhong , Yonghao Long , Qi Dou , Yun-Hui Liu , Dong Sun

Monocular 3D object detection aims to extract the 3D position and properties of objects from a 2D input image. This is an ill-posed problem with a major difficulty lying in the information loss by depth-agnostic cameras. Conventional…

计算机视觉与模式识别 · 计算机科学 2020-09-01 Lijie Liu , Chufan Wu , Jiwen Lu , Lingxi Xie , Jie Zhou , Qi Tian

Reconstructing 3D scenes from monocular surgical videos can enhance surgeon's perception and therefore plays a vital role in various computer-assisted surgery tasks. However, achieving scale-consistent reconstruction remains an open…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Jiaxin Guo , Wenzhen Dong , Tianyu Huang , Hao Ding , Ziyi Wang , Haomin Kuang , Qi Dou , Yun-Hui Liu

Building articulated objects is a key challenge in computer vision. Existing methods often fail to effectively integrate information across different object states, limiting the accuracy of part-mesh reconstruction and part dynamics…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Yu Liu , Baoxiong Jia , Ruijie Lu , Junfeng Ni , Song-Chun Zhu , Siyuan Huang

We present Mono-STAR, the first real-time 3D reconstruction system that simultaneously supports semantic fusion, fast motion tracking, non-rigid object deformation, and topological change under a unified framework. The proposed system…

机器人学 · 计算机科学 2023-02-01 Haonan Chang , Dhruv Metha Ramesh , Shijie Geng , Yuqiu Gan , Abdeslam Boularias

3D object detection from monocular images has proven to be an enormously challenging task, with the performance of leading systems not yet achieving even 10\% of that of LiDAR-based counterparts. One explanation for this performance gap is…

计算机视觉与模式识别 · 计算机科学 2018-11-21 Thomas Roddick , Alex Kendall , Roberto Cipolla

This paper focuses on the challenging problem of 3D pose estimation of a diverse spectrum of articulated objects from single depth images. A novel structured prediction approach is considered, where 3D poses are represented as skeletal…

计算机视觉与模式识别 · 计算机科学 2016-12-05 Yu Zhang , Chi Xu , Li Cheng

We present MOSAIC-GS, a novel, fully explicit, and computationally efficient approach for high-fidelity dynamic scene reconstruction from monocular videos using Gaussian Splatting. Monocular reconstruction is inherently ill-posed due to the…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Svitlana Morkva , Maximum Wilder-Smith , Michael Oechsle , Alessio Tonioni , Marco Hutter , Vaishakh Patil

Vision-Language Models (VLMs) have achieved strong performance on implicit and explicit visual grounding and related tasks. However, such abilities are generally tested on simple, single-object phrases. We find that grounding performance…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Jiayun Luo , Mir Rayat Imtiaz Hossain , Pritam Sarkar , Boyang Li , Leonid Sigal

The precise localization of 3D objects from a single image without depth information is a highly challenging problem. Most existing methods adopt the same approach for all objects regardless of their diverse distributions, leading to…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Yunpeng Zhang , Jiwen Lu , Jie Zhou

Recovering a textured 3D mesh from a monocular image is highly challenging, particularly for in-the-wild objects that lack 3D ground truths. In this work, we present MeshInversion, a novel framework to improve the reconstruction by…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Junzhe Zhang , Daxuan Ren , Zhongang Cai , Chai Kiat Yeo , Bo Dai , Chen Change Loy

Object pose estimation is a non-trivial task that enables robotic manipulation, bin picking, augmented reality, and scene understanding, to name a few use cases. Monocular object pose estimation gained considerable momentum with the rise of…

计算机视觉与模式识别 · 计算机科学 2023-07-24 Stefan Thalhammer , Peter Hönig , Jean-Baptiste Weibel , Markus Vincze

This paper presents a robust monocular visual SLAM system that simultaneously utilizes point, line, and vanishing point features for accurate camera pose estimation and mapping. To address the critical challenge of achieving reliable…

机器人学 · 计算机科学 2025-03-13 Bingzheng Jiang , Jiayuan Wang , Han Ding , Lijun Zhu

Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Lixin Xue , Chengwei Zheng , Georgios Paschalidis , Chen Guo , Manuel Kaufmann , Juan Zarate , Dimitrios Tzionas