中文
相关论文

相关论文: Hierarchical Open-Vocabulary 3D Scene Graphs for L…

200 篇论文

To complete a complex task where a robot navigates to a goal object and fetches it, the robot needs to have a good understanding of the instructions and the surrounding environment. Large pre-trained models have shown capabilities to…

机器人学 · 计算机科学 2024-08-21 Yu Li , Dayou Li , Chenkun Zhao , Ruifeng Wang , Ran Song , Wei Zhang

Visual spatial description (VSD) aims to generate texts that describe the spatial relations of the given objects within images. Existing VSD work merely models the 2D geometrical vision features, thus inevitably falling prey to the problem…

计算机视觉与模式识别 · 计算机科学 2023-05-26 Yu Zhao , Hao Fei , Wei Ji , Jianguo Wei , Meishan Zhang , Min Zhang , Tat-Seng Chua

3D Gaussian Splatting (3DGS) has become horsepower in high-quality, real-time rendering for novel view synthesis of 3D scenes. However, existing methods focus primarily on geometric and appearance modeling, lacking deeper scene…

图形学 · 计算机科学 2025-07-01 Minchao Jiang , Shunyu Jia , Jiaming Gu , Xiaoyuan Lu , Guangming Zhu , Anqi Dong , Liang Zhang

Mapping and localization are two essential tasks for mobile robots in real-world applications. However, largescale and dynamic scenes challenge the accuracy and robustness of most current mature solutions. This situation becomes even worse…

机器人学 · 计算机科学 2022-01-19 Fan Wang , Chaofan Zhang , Fulin Tang , Hongkui Jiang , Yihong Wu , Yong Liu

Generating dialogue grounded in videos requires a high level of understanding and reasoning about the visual scenes in the videos. However, existing large visual-language models are not effective due to their latent features and…

计算机视觉与模式识别 · 计算机科学 2023-11-23 Hongcheng Liu , Zhe Chen , Hui Li , Pingjie Wang , Yanfeng Wang , Yu Wang

Deep Learning has revolutionized our ability to solve complex problems such as Vision-and-Language Navigation (VLN). This task requires the agent to navigate to a goal purely based on visual sensory inputs given natural language…

机器人学 · 计算机科学 2021-04-22 Muhammad Zubair Irshad , Chih-Yao Ma , Zsolt Kira

Open-vocabulary 3D object detection for autonomous driving aims to detect novel objects beyond the predefined training label sets in point cloud scenes. Existing approaches achieve this by connecting traditional 3D object detectors with…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Adrian Chow , Evelien Riddell , Yimu Wang , Sean Sedwards , Krzysztof Czarnecki

Grounding open-ended semantic instructions into physically executable local goals is a fundamental challenge in human-robot interaction. While existing navigation frameworks often regress deterministic waypoints, this rigid formulation…

机器人学 · 计算机科学 2026-05-20 Kaijie Yun , Yue Chen

Fast, collision-free motion through unknown environments remains a challenging problem for robotic systems. In these situations, the robot's ability to reason about its future motion is often severely limited by sensor field of view (FOV).…

机器学习 · 计算机科学 2018-03-07 Kapil Katyal , Katie Popek , Chris Paxton , Joseph Moore , Kevin Wolfe , Philippe Burlina , Gregory D. Hager

We introduce ShelfGaussian, an open-vocabulary multi-modal Gaussian-based 3D scene understanding framework supervised by off-the-shelf vision foundation models (VFMs). Gaussian-based methods have demonstrated superior performance and…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Lingjun Zhao , Yandong Luo , James Hays , Lu Gan

Most Vision-and-Language Navigation (VLN) algorithms are prone to making inaccurate decisions due to their lack of visual common sense and limited reasoning capabilities. To address this issue, we propose a Hierarchical Spatial Proximity…

计算机视觉与模式识别 · 计算机科学 2024-10-23 Ming Xu , Zilong Xie

Traditional closed-set 3D detection frameworks fail to meet the demands of open-world applications like autonomous driving. Existing open-vocabulary 3D detection methods typically adopt a two-stage pipeline consisting of pseudo-label…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Qi Liu , Yabei Li , Hongsong Wang , Lei He

The 3D scene graph models spatial relationships between objects, enabling the agent to efficiently navigate in a partially observable environment and predict the location of the target object.This paper proposes an original framework named…

机器人学 · 计算机科学 2025-06-06 Nikita Oskolkov , Huzhenyu Zhang , Dmitry Makarov , Dmitry Yudin , Aleksandr Panov

Constructing a 3D scene capable of accommodating open-ended language queries, is a pivotal pursuit, particularly within the domain of robotics. Such technology facilitates robots in executing object manipulations based on human language…

Abstract semantic 3D scene understanding is a problem of critical importance in robotics. As robots still lack the common-sense knowledge about household objects and locations of an average human, we investigate the use of pre-trained…

机器人学 · 计算机科学 2023-11-09 William Chen , Siyi Hu , Rajat Talak , Luca Carlone

Connecting current observations with prior experiences helps robots adapt and plan in new, unseen 3D environments. Recently, 3D scene analogies have been proposed to connect two 3D scenes, which are smooth maps that align scene regions with…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Junho Kim , Young Min Kim

In this paper, we propose a new framework for zero-shot object navigation. Existing zero-shot object navigation methods prompt LLM with the text of spatially closed objects, which lacks enough scene context for in-depth reasoning. To better…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Hang Yin , Xiuwei Xu , Zhenyu Wu , Jie Zhou , Jiwen Lu

Scene graphs have emerged as a powerful tool for robots, providing a structured representation of spatial and semantic relationships for advanced task planning. Despite their potential, conventional 3D indoor scene graphs face critical…

机器人学 · 计算机科学 2025-10-17 Jeewon Kim , Minho Oh , Hyun Myung

Open-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions. However, complex visual scenes can be naturally decomposed into simpler parts and abstracted at multiple levels of…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Xudong Wang , Shufan Li , Konstantinos Kallidromitis , Yusuke Kato , Kazuki Kozuka , Trevor Darrell

3D open-vocabulary scene understanding, crucial for advancing augmented reality and robotic applications, involves interpreting and locating specific regions within a 3D space as directed by natural language instructions. To this end, we…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Yansong Qu , Shaohui Dai , Xinyang Li , Jianghang Lin , Liujuan Cao , Shengchuan Zhang , Rongrong Ji
‹ 上一页 1 8 9 10 下一页 ›