中文
相关论文

相关论文: Fin3R: Fine-tuning Feed-forward 3D Reconstruction …

200 篇论文

Monocular depth estimation, enabled by self-supervised learning, is a key technique for 3D perception in computer vision. However, it faces significant challenges in real-world scenarios, which encompass adverse weather variations, motion…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Runze Chen , Haiyong Luo , Fang Zhao , Jingze Yu , Yupeng Jia , Juan Wang , Xuepeng Ma

Recovering the 3D geometry of a scene from a sparse set of uncalibrated images is a long-standing problem in computer vision. While recent learning-based approaches such as DUSt3R and MASt3R have demonstrated impressive results by directly…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Sara Rojas , Matthieu Armando , Bernard Ghamen , Philippe Weinzaepfel , Vincent Leroy , Gregory Rogez

Fine-grained image retrieval (FGIR) is to learn visual representations that distinguish visually similar objects while maintaining generalization. Existing methods propose to generate discriminative features, but rarely consider the…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Xin Jiang , Hao Tang , Rui Yan , Jinhui Tang , Zechao Li

We present PAD3R, a method for reconstructing deformable 3D objects from casually captured, unposed monocular videos. Unlike existing approaches, PAD3R handles long video sequences featuring substantial object deformation, large-scale…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ting-Hsuan Liao , Haowen Liu , Yiran Xu , Songwei Ge , Gengshan Yang , Jia-Bin Huang

We introduce a novel task of 3D visual grounding in monocular RGB images using language descriptions with both appearance and geometry information. Specifically, we build a large-scale dataset, Mono3DRefer, which contains 3D object targets…

计算机视觉与模式识别 · 计算机科学 2023-12-14 Yang Zhan , Yuan Yuan , Zhitong Xiong

Single-image 3D reconstruction with large reconstruction models (LRMs) has advanced rapidly, yet reconstructions often exhibit geometric inconsistencies and misaligned details that limit fidelity. We introduce GeoFusionLRM, a geometry-aware…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Ahmet Burak Yildirim , Tuna Saygin , Duygu Ceylan , Aysegul Dundar

3D reassembly is a fundamental geometric problem, and in recent years it has increasingly been challenged by deep learning methods rather than classical optimization. While learning approaches have shown promising results, most still rely…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Adeela Islam , Stefano Fiorini , Manuel Lecha , Theodore Tsesmelis , Stuart James , Pietro Morerio , Alessio Del Bue

Accurate depth estimation is fundamental to 3D perception in autonomous driving, supporting tasks such as detection, tracking, and motion planning. However, monocular camera-based 3D detection suffers from depth ambiguity and reduced…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Chen-Chou Lo , Patrick Vandewalle

Despite significant progress made in the past few years, challenges remain for depth estimation using a single monocular image. First, it is nontrivial to train a metric-depth prediction model that can generalize well to diverse scenes…

计算机视觉与模式识别 · 计算机科学 2022-09-07 Wei Yin , Jianming Zhang , Oliver Wang , Simon Niklaus , Simon Chen , Yifan Liu , Chunhua Shen

The rapid development of Large Multimodal Models (LMMs) has led to remarkable progress in 2D visual understanding; however, extending these capabilities to 3D scene understanding remains a significant challenge. Existing approaches…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Hongpei Zheng , Lintao Xiang , Qijun Yang , Qian Lin , Hujun Yin

A reasonable and balanced diet is essential for maintaining good health. With the advancements in deep learning, automated nutrition estimation method based on food images offers a promising solution for monitoring daily nutritional intake…

计算机视觉与模式识别 · 计算机科学 2023-10-19 Yuzhe Han , Qimin Cheng , Wenjin Wu , Ziyang Huang

Fine-grained 3D shape retrieval aims to retrieve 3D shapes similar to a query shape in a repository with models belonging to the same class, which requires shape descriptors to be capable of representing detailed geometric information to…

计算机视觉与模式识别 · 计算机科学 2020-10-05 Rao Fu , Jie Yang , Jiawei Sun , Fang-Lue Zhang , Yu-Kun Lai , Lin Gao

Direct automatic segmentation of objects from 3D medical imaging, such as magnetic resonance (MR) imaging, is challenging as it often involves accurately identifying a number of individual objects with complex geometries within a large…

图像与视频处理 · 电气工程与系统科学 2021-09-23 Wei Dai , Boyeong Woo , Siyu Liu , Matthew Marques , Craig B. Engstrom , Peter B. Greer , Stuart Crozier , Jason A. Dowling , Shekhar S. Chandra

Reconstructing high-quality point clouds from images remains challenging in computer vision. Existing generative-model-based approaches, particularly diffusion-model approaches that directly learn the posterior, may suffer from…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Seunghyeok Shin , Dabin Kim , Hongki Lim

Reconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are built upon multi-view geometry. State-of-the-art (SOTA)…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Wei Yin , Chi Zhang , Hao Chen , Zhipeng Cai , Gang Yu , Kaixuan Wang , Xiaozhi Chen , Chunhua Shen

State-of-the-art techniques for monocular camera reconstruction predominantly rely on the Structure from Motion (SfM) pipeline. However, such methods often yield reconstruction outcomes that lack crucial scale information, and over time,…

机器人学 · 计算机科学 2023-10-10 Chunge Bai , Ruijie Fu , Xiang Gao

While recent feed-forward 3D reconstruction models provide a strong geometric foundation for scene understanding, extending them to 3D instance segmentation typically relies on a disjointed "lift-and-cluster" paradigm. Grouping dense…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Changyang Li , Xueqing Huang , Shin-Fang Chng , Huangying Zhan , Qingan Yan , Yi Xu

Volumetric models have become a popular representation for 3D scenes in recent years. One breakthrough leading to their popularity was KinectFusion, which focuses on 3D reconstruction using RGB-D sensors. However, monocular SLAM has since…

计算机视觉与模式识别 · 计算机科学 2017-08-03 Victor Adrian Prisacariu , Olaf Kähler , Stuart Golodetz , Michael Sapienza , Tommaso Cavallari , Philip H S Torr , David W Murray

We propose DiMeR, a novel geometry-texture disentangled feed-forward model with 3D supervision for sparse-view mesh reconstruction. Existing methods confront two persistent obstacles: (i) textures can conceal geometric errors, i.e.,…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Lutao Jiang , Jiantao Lin , Kanghao Chen , Wenhang Ge , Xin Yang , Yifan Jiang , Yuanhuiyi Lyu , Xu Zheng , Yinchuan Li , Yingcong Chen

In this paper, we propose enhancing monocular depth estimation by adding 3D points as depth guidance. Unlike existing depth completion methods, our approach performs well on extremely sparse and unevenly distributed point clouds, which…

计算机视觉与模式识别 · 计算机科学 2022-08-02 Lam Huynh , Phong Nguyen-Ha , Jiri Matas , Esa Rahtu , Janne Heikkila