中文
相关论文

相关论文: ScanRefer: 3D Object Localization in RGB-D Scans u…

200 篇论文

Although great progress has been made in 3D visual grounding, current models still rely on explicit textual descriptions for grounding and lack the ability to reason human intentions from implicit instructions. We propose a new task called…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Chenming Zhu , Tai Wang , Wenwei Zhang , Kai Chen , Xihui Liu

Dense captioning in 3D point clouds is an emerging vision-and-language task involving object-level 3D scene understanding. Apart from coarse semantic class prediction and bounding box regression as in traditional 3D object detection, 3D…

计算机视觉与模式识别 · 计算机科学 2022-04-25 Heng Wang , Chaoyi Zhang , Jianhui Yu , Weidong Cai

Natural language interfaces to embodied AI are becoming more ubiquitous in our daily lives. This opens up further opportunities for language-based interaction with embodied agents, such as a user verbally instructing an agent to execute…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Jiaqi Chen , Daniel Barath , Iro Armeni , Marc Pollefeys , Hermann Blum

We propose a system that learns to detect objects and infer their 3D poses in RGB-D images. Many existing systems can identify objects and infer 3D poses, but they heavily rely on human labels and 3D annotations. The challenge here is to…

计算机视觉与模式识别 · 计算机科学 2020-11-02 Mihir Prabhudesai , Shamit Lal , Hsiao-Yu Fish Tung , Adam W. Harley , Shubhankar Potdar , Katerina Fragkiadaki

We propose an end-to-end solution to address the problem of object localisation in partial scenes, where we aim to estimate the position of an object in an unknown area given only a partial 3D scan of the scene. We propose a novel scene…

计算机视觉与模式识别 · 计算机科学 2022-11-02 Francesco Giuliari , Geri Skenderi , Marco Cristani , Alessio Del Bue , Yiming Wang

3D visual grounding aims to localize the object in 3D point cloud scenes that semantically corresponds to given natural language sentences. It is very critical for roadside infrastructure system to interpret natural languages and localize…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Panquan Yang , Junfei Huang , Zongzhangbao Yin , Yingsong Hu , Anni Xu , Xinyi Luo , Xueqi Sun , Hai Wu , Sheng Ao , Zhaoxing Zhu , Chenglu Wen , Cheng Wang

Recent advancements in 3D Large Language Models (LLMs) have demonstrated promising capabilities for 3D scene understanding. However, previous methods exhibit deficiencies in general referencing and grounding capabilities for intricate scene…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Haifeng Huang , Yilun Chen , Zehan Wang , Rongjie Huang , Runsen Xu , Tai Wang , Luping Liu , Xize Cheng , Yang Zhao , Jiangmiao Pang , Zhou Zhao

We propose an end-to-end approach to the natural language object retrieval task, which localizes an object within an image according to a natural language description, i.e., referring expression. Previous works divide this problem into two…

计算机视觉与模式识别 · 计算机科学 2017-03-23 Fan Wu , Zhongwen Xu , Yi Yang

Advances in deep learning techniques have allowed recent work to reconstruct the shape of a single object given only one RBG image as input. Building on common encoder-decoder architectures for this task, we propose three extensions: (1)…

计算机视觉与模式识别 · 计算机科学 2020-08-06 Stefan Popov , Pablo Bauszat , Vittorio Ferrari

3D visual grounding is the task of localizing the object in a 3D scene which is referred by a description in natural language. With a wide range of applications ranging from autonomous indoor robotics to AR/VR, the task has recently risen…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Ozan Unal , Christos Sakaridis , Suman Saha , Luc Van Gool

We present a novel task for cross-dataset visual grounding in 3D scenes (Cross3DVG), which overcomes limitations of existing 3D visual grounding models, specifically their restricted 3D resources and consequent tendencies of overfitting a…

计算机视觉与模式识别 · 计算机科学 2024-02-08 Taiki Miyanishi , Daichi Azuma , Shuhei Kurita , Motoki Kawanabe

Designing 3D scenes is currently a creative task that requires significant expertise and effort in using complex 3D design interfaces. This effortful design process starts in stark contrast to the easiness with which people can use language…

图形学 · 计算机科学 2017-03-02 Angel X. Chang , Mihail Eric , Manolis Savva , Christopher D. Manning

Humans describe the physical world using natural language to refer to specific 3D locations based on a vast range of properties: visual appearance, semantics, abstract associations, or actionable affordances. In this work we propose…

计算机视觉与模式识别 · 计算机科学 2023-03-17 Justin Kerr , Chung Min Kim , Ken Goldberg , Angjoo Kanazawa , Matthew Tancik

We propose Scan2Part, a method to segment individual parts of objects in real-world, noisy indoor RGB-D scans. To this end, we vary the part hierarchies of objects in indoor scenes and explore their effect on scene understanding models.…

计算机视觉与模式识别 · 计算机科学 2022-06-07 Alexandr Notchenko , Vladislav Ishimtsev , Alexey Artemov , Vadim Selyutin , Emil Bogomolov , Evgeny Burnaev

The 3D visual grounding task has been explored with visual and language streams comprehending referential language to identify target objects in 3D scenes. However, most existing methods devote the visual stream to capturing the 3D visual…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Eslam Mohamed Bakr , Yasmeen Alsaedy , Mohamed Elhoseiny

In this work, we introduce the task of 3D object instance re-localization (RIO): given one or multiple objects in an RGB-D scan, we want to estimate their corresponding 6DoF poses in another 3D scan of the same environment taken at a later…

计算机视觉与模式识别 · 计算机科学 2019-08-20 Johanna Wald , Armen Avetisyan , Nassir Navab , Federico Tombari , Matthias Nießner

Recent progress in 3D scene understanding has explored visual grounding (3DVG) to localize a target object through a language description. However, existing methods only consider the dependency between the entire sentence and the target…

计算机视觉与模式识别 · 计算机科学 2023-05-30 Zhihao Yuan , Xu Yan , Zhuo Li , Xuhao Li , Yao Guo , Shuguang Cui , Zhen Li

We present a 3D object detection method that uses regressed descriptors of locally-sampled RGB-D patches for 6D vote casting. For regression, we employ a convolutional auto-encoder that has been trained on a large collection of random local…

计算机视觉与模式识别 · 计算机科学 2016-07-21 Wadim Kehl , Fausto Milletari , Federico Tombari , Slobodan Ilic , Nassir Navab

We propose associating language utterances to 3D visual abstractions of the scene they describe. The 3D visual abstractions are encoded as 3-dimensional visual feature maps. We infer these 3D visual scene feature maps from RGB images of the…

计算机视觉与模式识别 · 计算机科学 2021-06-21 Mihir Prabhudesai , Hsiao-Yu Fish Tung , Syed Ashar Javed , Maximilian Sieb , Adam W. Harley , Katerina Fragkiadaki

One major goal of vision is to infer physical models of objects, surfaces, and their layout from sensors. In this paper, we aim to interpret indoor scenes from one RGBD image. Our representation encodes the layout of orthogonal walls and…

计算机视觉与模式识别 · 计算机科学 2018-11-15 Chuhang Zou , Ruiqi Guo , Zhizhong Li , Derek Hoiem