中文
相关论文

相关论文: Fin3R: Fine-tuning Feed-forward 3D Reconstruction …

200 篇论文

Monocular image-based 3D reconstruction of faces is a long-standing problem in computer vision. Since image data is a 2D projection of a 3D face, the resulting depth ambiguity makes the problem ill-posed. Most existing methods rely on…

Monocular 3D face reconstruction plays a crucial role in avatar generation, with significant demand in web-related applications such as generating virtual financial advisors in FinTech. Current reconstruction methods predominantly rely on…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Haoxin Xu , Zezheng Zhao , Yuxin Cao , Chunyu Chen , Hao Ge , Ziyao Liu

This paper investigates the research task of reconstructing the 3D clothed human body from a monocular image. Due to the inherent ambiguity of single-view input, existing approaches leverage pre-trained SMPL(-X) estimation models or…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Gangjian Zhang , Nanjie Yao , Shunsi Zhang , Hanfeng Zhao , Guoliang Pang , Jian Shu , Hao Wang

Any entity in the visual world can be hierarchically grouped based on shared characteristics and mapped to fine-grained sub-categories. While Multi-modal Large Language Models (MLLMs) achieve strong performance on coarse-grained visual…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Hulingxiao He , Zijun Geng , Yuxin Peng

Monocular 3D foundation models offer an extensible solution for perception tasks, making them attractive for broader 3D vision applications. In this paper, we propose MoRe, a training-free Monocular Geometry Refinement method designed to…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Dongki Jung , Jaehoon Choi , Yonghan Lee , Sungmin Eum , Heesung Kwon , Dinesh Manocha

The precise reconstruction of 3D objects from a single RGB image in complex scenes presents a critical challenge in virtual reality, autonomous driving, and robotics. Existing neural implicit 3D representation methods face significant…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Luoxi Zhang , Pragyan Shrestha , Yu Zhou , Chun Xie , Itaru Kitahara

Feed-forward 3D reconstruction methods aim to predict the 3D structure of a scene directly from input images, providing a faster alternative to per-scene optimization approaches. Significant progress has been made in single-view and…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Sam Bahrami , Dylan Campbell

Large Foundation Models like Dust3r can produce high quality outputs such as pointmaps, camera intrinsics, and depth estimation, given stereo-image pairs as input. However, the application of these outputs on tasks like Visual Localization…

计算机视觉与模式识别 · 计算机科学 2026-02-20 Aditya Dutt , Ishikaa Lunawat , Manpreet Kaur

Recent advancements in neural visual geometry, including transformer-based models such as VGGT and Pi3, have achieved impressive accuracy on 3D reconstruction tasks. However, their reliance on full attention makes them fundamentally limited…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Leo Kaixuan Cheng , Abdus Shaikh , Ruofan Liang , Zhijie Wu , Yushi Guan , Nandita Vijaykumar

Multi-view 3D reconstruction has remained an essential yet challenging problem in the field of computer vision. While DUSt3R and its successors have achieved breakthroughs in 3D reconstruction from unposed images, these methods exhibit…

图像与视频处理 · 电气工程与系统科学 2025-09-16 Sidun Liu , Wenyu Li , Peng Qiao , Yong Dou

Object-centric scene understanding is a fundamental challenge in computer vision. Existing approaches often rely on multi-stage pipelines that first apply pre-trained segmentors to extract individual objects, followed by per-object 3D…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Yi Du , Yang You , Xiang Wan , Leonidas Guibas

Federated Learning (FL) methods often struggle in highly statistically heterogeneous settings. Indeed, non-IID data distributions cause client drift and biased local solutions, particularly pronounced in the final classification layer,…

机器学习 · 计算机科学 2024-06-04 Eros Fanì , Raffaello Camoriano , Barbara Caputo , Marco Ciccone

Current methods for dense 3D point tracking in dynamic scenes typically rely on pairwise processing, require known camera poses, or assume temporal ordering of input frames, thereby constraining their flexibility and applicability.…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Vivek Alumootil , Tuan-Anh Vu

We present MoGe, a powerful model for recovering 3D geometry from monocular open-domain images. Given a single image, our model directly predicts a 3D point map of the captured scene with an affine-invariant representation, which is…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Ruicheng Wang , Sicheng Xu , Cassie Dai , Jianfeng Xiang , Yu Deng , Xin Tong , Jiaolong Yang

Panoptic segmentation of 3D scenes, involving the segmentation and classification of object instances in a dense 3D reconstruction of a scene, is a challenging problem, especially when relying solely on unposed 2D images. Existing…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Lojze Zust , Yohann Cabon , Juliette Marrie , Leonid Antsfeld , Boris Chidlovskii , Jerome Revaud , Gabriela Csurka

We propose MoGe-2, an advanced open-domain geometry estimation model that recovers a metric scale 3D point map of a scene from a single image. Our method builds upon the recent monocular geometry estimation approach, MoGe, which predicts…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Ruicheng Wang , Sicheng Xu , Yue Dong , Yu Deng , Jianfeng Xiang , Zelong Lv , Guangzhong Sun , Xin Tong , Jiaolong Yang

Recently, neural implicit 3D reconstruction in indoor scenarios has become popular due to its simplicity and impressive performance. Previous works could produce complete results leveraging monocular priors of normal or depth. However, they…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Xinghui Li , Yuchen Ji , Xiansong Lai , Wanting Zhang

Composed Image Retrieval (CIR) facilitates image retrieval through a multimodal query consisting of a reference image and modification text. The reference image defines the retrieval context, while the modification text specifies desired…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Zixu Li , Zhiheng Fu , Yupeng Hu , Zhiwei Chen , Haokun Wen , Liqiang Nie

We present a unified framework capable of solving a broad range of 3D tasks. Our approach features a stateful recurrent model that continuously updates its state representation with each new observation. Given a stream of images, this…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Qianqian Wang , Yifei Zhang , Aleksander Holynski , Alexei A. Efros , Angjoo Kanazawa

Visual SLAM is a cornerstone technique in robotics, autonomous driving and extended reality (XR), yet classical systems often struggle with low-texture environments, scale ambiguity, and degraded performance under challenging visual…

机器人学 · 计算机科学 2025-11-18 Yuxuan Zhou , Xingxing Li , Shengyu Li , Zhuohao Yan , Chunxi Xia , Shaoquan Feng