English
Related papers

Related papers: DOZE: A Dataset for Open-Vocabulary Zero-Shot Obje…

200 papers

Traditional 3D scene understanding approaches rely on labeled 3D datasets to train a model for a single task with supervision. We propose OpenScene, an alternative approach where a model predicts dense features for 3D scene points that are…

Computer Vision and Pattern Recognition · Computer Science 2023-04-07 Songyou Peng , Kyle Genova , Chiyu "Max" Jiang , Andrea Tagliasacchi , Marc Pollefeys , Thomas Funkhouser

Object Goal Navigation (ObjectNav) challenges robots to find objects in unseen environments, demanding sophisticated reasoning. While Vision-Language Models (VLMs) show potential, current ObjectNav methods often employ them superficially,…

Robotics · Computer Science 2025-06-23 Mobin Habibpour , Fatemeh Afghah

Scene understanding and reasoning has been a fundamental problem in 3D computer vision, requiring models to identify objects, their properties, and spatial or comparative relationships among the objects. Existing approaches enable this by…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Vivek Madhavaram , Vartika Sengar , Arkadipta De , Charu Sharma

We present LGX (Language-guided Exploration), a novel algorithm for Language-Driven Zero-Shot Object Goal Navigation (L-ZSON), where an embodied agent navigates to a uniquely described target object in a previously unseen environment. Our…

Robotics · Computer Science 2024-04-16 Vishnu Sashank Dorbala , James F. Mullen , Dinesh Manocha

The ability to accurately locate and navigate to a specific object is a crucial capability for embodied agents that operate in the real world and interact with objects to complete tasks. Such object navigation tasks usually require…

Artificial Intelligence · Computer Science 2023-07-07 Kaiwen Zhou , Kaizhi Zheng , Connor Pryor , Yilin Shen , Hongxia Jin , Lise Getoor , Xin Eric Wang

Building a general-purpose intelligent home-assistant agent skilled in diverse tasks by human commands is a long-term blueprint of embodied AI research, which poses requirements on task planning, environment modeling, and object…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Xinyu Xu , Shengcheng Luo , Yanchao Yang , Yong-Lu Li , Cewu Lu

Visual Language Navigation is a task that challenges robots to navigate in realistic environments based on natural language instructions. While previous research has largely focused on static settings, real-world navigation must often…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Dillon Loh , Tomasz Bednarz , Xinxing Xia , Frank Guan

Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation for embodied intelligence. For UAV agents in particular, accurate prediction of complex…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Zile Guo , Zhan Chen , Enze Zhu , Kan Wei , Yongkang Zou , Xiaoxuan Liu , Lei Wang

This work targets what we consider to be the foundational step for urban airborne robots, a safe landing. Our attention is directed toward what we deem the most crucial aspect of the safe landing perception stack: segmentation. We present a…

Robotics · Computer Science 2024-10-16 Haechan Mark Bong , Rongge Zhang , Ricardo de Azambuja , Giovanni Beltrame

Socially compliant navigation requires structured reasoning over dynamic pedestrians and physical constraints to ensure safe and interpretable decisions. However, existing social navigation datasets often lack explicit reasoning supervision…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Zhuonan Liu , Xinyu Zhang , Zishuo Wang , Tomohito Kawabata , Xuesu Xiao , Ling Xiao

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

Computer Vision and Pattern Recognition · Computer Science 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

Object goal navigation is a fundamental task in embodied AI, where an agent is instructed to locate a target object in an unexplored environment. Traditional learning-based methods rely heavily on large-scale annotated data or require…

Robotics · Computer Science 2025-06-05 Arnab Debnath , Gregory J. Stein , Jana Kosecka

What is a good visual representation for autonomous agents? We address this question in the context of semantic visual navigation, which is the problem of a robot finding its way through a complex environment to a target object, e.g. go to…

Computer Vision and Pattern Recognition · Computer Science 2019-07-04 Arsalan Mousavian , Alexander Toshev , Marek Fiser , Jana Kosecka , Ayzaan Wahid , James Davidson

Estimating the 6D pose of objects unseen during training is highly desirable yet challenging. Zero-shot object 6D pose estimation methods address this challenge by leveraging additional task-specific supervision provided by large-scale,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Andrea Caraffa , Davide Boscaini , Amir Hamza , Fabio Poiesi

In the realm of object pose estimation, scenarios involving both dynamic objects and moving cameras are prevalent. However, the scarcity of corresponding real-world datasets significantly hinders the development and evaluation of robust…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Xiangting Meng , Jiaqi Yang , Mingshu Chen , Chenxin Yan , Yujiao Shi , Wenchao Ding , Laurent Kneip

Zero-shot detection, namely, localizing both seen and unseen objects, increasingly gains importance for large-scale applications, with large number of object classes, since, collecting sufficient annotated data with ground truth bounding…

Computer Vision and Pattern Recognition · Computer Science 2020-04-13 Pengkai Zhu , Hanxiao Wang , Venkatesh Saligrama

Current perception models in autonomous driving have become notorious for greatly relying on a mass of annotated data to cover unseen cases and address the long-tail problem. On the other hand, learning from unlabeled large-scale collected…

Computer Vision and Pattern Recognition · Computer Science 2021-10-26 Jiageng Mao , Minzhe Niu , Chenhan Jiang , Hanxue Liang , Jingheng Chen , Xiaodan Liang , Yamin Li , Chaoqiang Ye , Wei Zhang , Zhenguo Li , Jie Yu , Hang Xu , Chunjing Xu

Significant progress has been made in open-vocabulary mobile manipulation, where the goal is for a robot to perform tasks in any environment given a natural language description. However, most current systems assume a static environment,…

In driver activity monitoring, movements are mostly limited to the upper body, which makes many actions look similar. To tell these actions apart, human often rely on the objects the driver is using, such as holding a phone compared with…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Yiming Li , Chen Cai , Tianyi Liu , Dan Lin , Wenqian Wang , Wenfei Liang , Bingbing Li , Kim-Hui Yap

In the real world, autonomous driving agents navigate in highly dynamic environments full of unexpected situations where pre-trained models are unreliable. In these situations, what is immediately available to vehicles is often only human…

Artificial Intelligence · Computer Science 2022-10-25 Ziqiao Ma , Ben VanDerPloeg , Cristian-Paul Bara , Huang Yidong , Eui-In Kim , Felix Gervits , Matthew Marge , Joyce Chai