中文
相关论文

相关论文: Perspective from a Higher Dimension: Can 3D Geomet…

200 篇论文

In computer vision, camera pose estimation from correspondences between 3D geometric entities and their projections into the image has been a widely investigated problem. Although most state-of-the-art methods exploit low-level primitives…

计算机视觉与模式识别 · 计算机科学 2023-06-16 Vincent Gaudillière , Gilles Simon , Marie-Odile Berger

Visual object counting is a fundamental computer vision task underpinning numerous real-world applications, from cell counting in biomedicine to traffic and wildlife monitoring. However, existing methods struggle to handle the challenge of…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Corentin Dumery , Noa Etté , Aoxiang Fan , Ren Li , Jingyi Xu , Hieu Le , Pascal Fua

3D object detection is an indispensable component for scene understanding. However, the annotation of large-scale 3D datasets requires significant human effort. To tackle this problem, many methods adopt weakly supervised 3D object…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Guowen Zhang , Junsong Fan , Liyi Chen , Zhaoxiang Zhang , Zhen Lei , Lei Zhang

Visual localization is the problem of estimating the camera pose of a given query image within a known scene. Most state-of-the-art localization approaches follow the structure-based paradigm and use 2D-3D matches between pixels in a query…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Vojtech Panek , Torsten Sattler , Zuzana Kukelova

Recent studies have shown remarkable advances in 3D human pose estimation from monocular images, with the help of large-scale in-door 3D datasets and sophisticated network architectures. However, the generalizability to different…

计算机视觉与模式识别 · 计算机科学 2019-03-28 Xipeng Chen , Kwan-Yee Lin , Wentao Liu , Chen Qian , Xiaogang Wang , Liang Lin

Manipulation planning is the problem of finding a sequence of robot configurations that involves interactions with objects in the scene, e.g., grasping and placing an object, or more general tool-use. To achieve such interactions,…

机器人学 · 计算机科学 2022-08-01 Jung-Su Ha , Danny Driess , Marc Toussaint

Recent developments in Multimodal Large Language Models (MLLMs) have significantly improved Vision-Language (VL) reasoning in 2D domains. However, extending these capabilities to 3D scene understanding remains a major challenge. Existing 3D…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Haijier Chen , Bo Xu , Shoujian Zhang , Haoze Liu , Jiaxuan Lin , Jingrong Wang

We propose SelfSplat, a novel 3D Gaussian Splatting model designed to perform pose-free and 3D prior-free generalizable 3D reconstruction from unposed multi-view images. These settings are inherently ill-posed due to the lack of…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Gyeongjin Kang , Jisang Yoo , Jihyeon Park , Seungtae Nam , Hyeonsoo Im , Sangheon Shin , Sangpil Kim , Eunbyung Park

The 3D depth estimation and relative pose estimation problem within a decentralized architecture is a challenging problem that arises in missions that require coordination among multiple vision-controlled robots. The depth estimation…

机器人学 · 计算机科学 2019-08-02 Romulo T. Rodrigues , Pedro Miraldo , Dimos V. Dimarogonas , A. Pedro Aguiar

We propose an active learning approach to image segmentation that exploits geometric priors to speed up and streamline the annotation process. It can be applied for both background-foreground and multi-class segmentation tasks in 2D images…

计算机视觉与模式识别 · 计算机科学 2019-04-05 Ksenia Konyushkova , Raphael Sznitman , Pascal Fua

We introduce a test-time framework for multiview Transformers (MVTs) that incorporates priors (e.g., camera poses, intrinsics, and depth) to improve 3D tasks without retraining or modifying pre-trained image-only networks. Rather than…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Lei Zhou , Haoyu Wu , Akshat Dave , Dimitris Samaras

Hyperbolic spaces allow for more efficient modeling of complex, hierarchical structures, which is particularly beneficial in tasks involving multi-modal data. Although hyperbolic geometries have been proven effective for language-image…

计算机视觉与模式识别 · 计算机科学 2025-01-08 Yingjie Liu , Pengyu Zhang , Ziyao He , Mingsong Chen , Xuan Tang , Xian Wei

Online real estate platforms have become significant marketplaces facilitating users' search for an apartment or a house. Yet it remains challenging to accurately appraise a property's value. Prior works have primarily studied real estate…

机器学习 · 计算机科学 2021-02-17 Kirill Solovev , Nicolas Pröllochs

The large variation of viewpoint and irrelevant content around the target always hinder accurate image retrieval and its subsequent tasks. In this paper, we investigate an extremely challenging task: given a ground-view image of a landmark,…

计算机视觉与模式识别 · 计算机科学 2022-05-24 Zelong Zeng , Zheng Wang , Fan Yang , Shin'ichi Satoh

Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent studies represent 3D scenes as 3D spatial videos composed of image sequences with depth…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Han Li , Zehao Huang , Jiahui Fu , Naiyan Wang , Si Liu

Visual (re)localization is critical for various applications in computer vision and robotics. Its goal is to estimate the 6 degrees of freedom (DoF) camera pose for each query image, based on a set of posed database images. Currently, all…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Siyan Dong , Shaohui Liu , Hengkai Guo , Baoquan Chen , Marc Pollefeys

Object-level Simultaneous Localization and Mapping (SLAM), which incorporates semantic information for high-level scene understanding, faces challenges of under-constrained optimization due to sparse observations. Prior work has introduced…

机器人学 · 计算机科学 2025-09-29 Yang Jiao , Yiding Qiu , Henrik I. Christensen

Realistic 3D indoor scene datasets have enabled significant recent progress in computer vision, scene understanding, autonomous navigation, and 3D reconstruction. But the scale, diversity, and customizability of existing datasets is…

计算机视觉与模式识别 · 计算机科学 2021-12-13 Kai Wang , Xianghao Xu , Leon Lei , Selena Ling , Natalie Lindsay , Angel X. Chang , Manolis Savva , Daniel Ritchie

Visual localization is the task of estimating a camera pose in a known environment. In this paper, we utilize 3D Gaussian Splatting (3DGS)-based representations for accurate and privacy-preserving visual localization. We propose Gaussian…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Maxime Pietrantoni , Gabriela Csurka , Torsten Sattler

We present LaLaLoc to localise in environments without the need for prior visitation, and in a manner that is robust to large changes in scene appearance, such as a full rearrangement of furniture. Specifically, LaLaLoc performs…

计算机视觉与模式识别 · 计算机科学 2021-10-13 Henry Howard-Jenkins , Jose-Raul Ruiz-Sarmiento , Victor Adrian Prisacariu