中文
相关论文

相关论文: Category-Level 3D Correspondence in Camera Space v…

200 篇论文

Finding correspondences is a fundamental and extensively researched problem in computer vision and graphics. In this work, we examine the underexplored task of estimating segmentation-to-segmentation correspondence between images in the…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Itai Lang , Dongwei Lyu , Dale Decatur , Rana Hanocka

Towards holistic understanding of 3D scenes, a general 3D segmentation method is needed that can segment diverse objects without restrictions on object quantity or categories, while also reflecting the inherent hierarchical structure. To…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Haiyang Ying , Yixuan Yin , Jinzhi Zhang , Fan Wang , Tao Yu , Ruqi Huang , Lu Fang

Object recognition has seen significant progress in the image domain, with focus primarily on 2D perception. We propose to leverage existing large-scale datasets of 3D models to understand the underlying 3D structure of objects seen in an…

计算机视觉与模式识别 · 计算机科学 2020-07-28 Weicheng Kuo , Anelia Angelova , Tsung-Yi Lin , Angela Dai

As a crucial task of autonomous driving, 3D object detection has made great progress in recent years. However, monocular 3D object detection remains a challenging problem due to the unsatisfactory performance in depth estimation. Most…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Yinmin Zhang , Xinzhu Ma , Shuai Yi , Jun Hou , Zhihui Wang , Wanli Ouyang , Dan Xu

Multi-image spatial reasoning remains challenging for current multimodal large language models (MLLMs). While single-view perception is inherently 2D, reasoning over multiple views requires building a coherent scene understanding across…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Xuejun Zhang , Aditi Tiwari , Zhenhailong Wang , Heng Ji

Given a single scene image, this paper proposes a method of Category-level 6D Object Pose and Size Estimation (COPSE) from the point cloud of the target object, without external real pose-annotated training data. Specifically, beyond the…

计算机视觉与模式识别 · 计算机科学 2022-04-12 Haitao Lin , Zichang Liu , Chilam Cheang , Yanwei Fu , Guodong Guo , Xiangyang Xue

Most deep learning approaches to comprehensive semantic modeling of 3D indoor spaces require costly dense annotations in the 3D domain. In this work, we explore a central 3D scene modeling task, namely, semantic scene reconstruction without…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Junwen Huang , Alexey Artemov , Yujin Chen , Shuaifeng Zhi , Kai Xu , Matthias Nießner

Roadside monocular 3D detection requires detecting objects of predefined classes in an RGB frame and predicting their 3D attributes, such as bird's-eye-view (BEV) locations. It has broad applications in traffic control, vehicle-vehicle…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Yechi Ma , Yanan Li , Wei Hua , Shu Kong

We present a novel meta-learning approach for 6D pose estimation on unknown objects. In contrast to ``instance-level" and ``category-level" pose estimation methods, our algorithm learns object representation in a category-agnostic way,…

计算机视觉与模式识别 · 计算机科学 2023-10-20 Yumeng Li , Ning Gao , Hanna Ziesche , Gerhard Neumann

Articulated objects exist widely in the real world. However, previous 3D generative methods for unsupervised part decomposition are unsuitable for such objects, because they assume a spatially fixed part location, resulting in inconsistent…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Yuki Kawana , Yusuke Mukuta , Tatsuya Harada

We tackle the task of scalable unsupervised object-centric representation learning on 3D scenes. Existing approaches to object-centric representation learning show limitations in generalizing to larger scenes as their learning processes…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Tianyu Wang , Kee Siong Ng , Miaomiao Liu

Point clouds are a set of data points in space to represent the 3D geometry of objects. A fundamental step in the processing is to identify a subset of points to represent the shape. While traditional sampling methods often ignore to…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Pierre Onghena , Santiago Velasco-Forero , Beatriz Marcotegui

Multimodal language models (MLLMs) are increasingly being applied in real-world environments, necessitating their ability to interpret 3D spaces and comprehend temporal dynamics. Current methods often rely on specialized architectural…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Benlin Liu , Yuhao Dong , Yiqin Wang , Zixian Ma , Yansong Tang , Luming Tang , Yongming Rao , Wei-Chiu Ma , Ranjay Krishna

Object concepts play a foundational role in human visual cognition, enabling perception, memory, and interaction in the physical world. Inspired by findings in developmental neuroscience - where infants are shown to acquire object…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Haoqian Liang , Xiaohui Wang , Zhichao Li , Ya Yang , Naiyan Wang

Recent advances in large multimodal models suggest that explicit reasoning mechanisms play a critical role in improving model reliability, interpretability, and cross-modal alignment. While such reasoning-centric approaches have been proven…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Tianjiao Yu , Xinzhuo Li , Yifan Shen , Yuanzhe Liu , Ismini Lourentzou

Given two images, we can estimate the relative camera pose between them by establishing image-to-image correspondences. Usually, correspondences are 2D-to-2D and the pose we estimate is defined only up to scale. Some applications, aiming at…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Axel Barroso-Laguna , Sowmya Munukutla , Victor Adrian Prisacariu , Eric Brachmann

Monocular 3D object detection aims to detect objects in a 3D physical world from a single camera. However, recent approaches either rely on expensive LiDAR devices, or resort to dense pixel-wise depth estimation that causes prohibitive…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Wentao Bao , Qi Yu , Yu Kong

The 3D Morphable Model (3DMM) is a powerful statistical tool for representing 3D face shapes. To build a 3DMM, a training set of face scans in full point-to-point correspondence is required, and its modeling capabilities directly depend on…

计算机视觉与模式识别 · 计算机科学 2021-06-25 Claudio Ferrari , Stefano Berretti , Pietro Pala , Alberto Del Bimbo

Robotic manipulation systems operating in complex environments rely on perception systems that provide information about the geometry (pose and 3D shape) of the objects in the scene along with other semantic information such as object…

机器人学 · 计算机科学 2023-05-17 Shubham Agrawal , Nikhil Chavan-Dafle , Isaac Kasahara , Selim Engin , Jinwook Huh , Volkan Isler

Recent advances in large vision-language models (VLMs) have shown significant promise for 3D scene understanding. Existing VLM-based approaches typically align 3D scene features with the VLM's embedding space. However, this implicit…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Chen Li , Eric Peh , Basura Fernando