English
Related papers

Related papers: The Robotic Vision Scene Understanding Challenge

200 papers

General scene understanding for robotics requires flexible semantic representation, so that novel objects and structures which may not have been known at training time can be identified, segmented and grouped. We present an algorithm which…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Kirill Mazur , Edgar Sucar , Andrew J. Davison

In this work we present a novel approach to joint semantic localisation and scene understanding. Our work is motivated by the need for localisation algorithms which not only predict 6-DoF camera pose but also simultaneously recognise…

Computer Vision and Pattern Recognition · Computer Science 2019-09-24 Ignas Budvytis , Marvin Teichmann , Tomas Vojir , Roberto Cipolla

Significant progress has been made in scene understanding which seeks to build 3D, metric and object-oriented representations of the world. Concurrently, reinforcement learning has made impressive strides largely enabled by advances in…

Robotics · Computer Science 2020-11-23 Zachary Ravichandran , J. Daniel Griffith , Benjamin Smith , Costas Frost

We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model. Most previous 3D segmentation approaches are…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Wei Xu , Chunsheng Shi , Sifan Tu , Xin Zhou , Dingkang Liang , Xiang Bai

Current exploration methods struggle to search for shops or restaurants in unknown open-world environments due to the lack of prior knowledge. Humans can leverage venue maps that offer valuable scene priors to aid exploration planning by…

Robotics · Computer Science 2025-03-04 Chang Chen , Liang Lu , Lei Yang , Yinqiang Zhang , Yizhou Chen , Ruixing Jia , Jia Pan

We live in a 3D world, performing activities and interacting with objects in the indoor environments everyday. Indoor scenes are the most familiar and essential environments in everyone's life. In the virtual world, 3D indoor scenes are…

Graphics · Computer Science 2017-06-30 Rui Ma

Semantic scene segmentation plays a critical role in a wide range of robotics applications, e.g., autonomous navigation. These applications are accompanied by specific computational restrictions, e.g., operation on low-power GPUs, at…

Computer Vision and Pattern Recognition · Computer Science 2021-08-26 Maria Tzelepi , Anastasios Tefas

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

Computer Vision and Pattern Recognition · Computer Science 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

Scene-understanding is an important topic in the area of Computer Vision, and illustrates computational challenges with applications to a wide range of domains including remote sensing, surveillance, smart agriculture, robotics, autonomous…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Zachary A Daniels , Dimitris Metaxas

Recent trends in image understanding have pushed for holistic scene understanding models that jointly reason about various tasks such as object detection, scene recognition, shape analysis, contextual reasoning, and local appearance based…

Computer Vision and Pattern Recognition · Computer Science 2014-06-17 Roozbeh Mottaghi , Sanja Fidler , Alan Yuille , Raquel Urtasun , Devi Parikh

We introduce the first approach to solve the challenging problem of unsupervised 4D visual scene understanding for complex dynamic scenes with multiple interacting people from multi-view video. Our approach simultaneously estimates a…

Computer Vision and Pattern Recognition · Computer Science 2019-07-24 Armin Mustafa , Chris Russell , Adrian Hilton

Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have primarily focused on indoor household tasks, they face two…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Penglei Sun , Yaoxian Song , Xiangru Zhu , Xiang Liu , Qiang Wang , Yue Liu , Changqun Xia , Tiefeng Li , Yang Yang , Xiaowen Chu

Visual and scalar-field (e.g., chemical) sensing are two of the options robot teams can use to perceive their environments when performing tasks. We give the first comparison of the computational characteristic of visual and scalar-field…

Multiagent Systems · Computer Science 2022-05-11 Todd Wareham , Andrew Vardy

Object pose estimation is a core perception task that enables, for example, object grasping and scene understanding. The widely available, inexpensive and high-resolution RGB sensors and CNNs that allow for fast inference based on this…

Careful robot manipulation in every-day cluttered environments requires an accurate understanding of the 3D scene, in order to grasp and place objects stably and reliably and to avoid colliding with other objects. In general, we must…

Robotics · Computer Science 2025-11-11 Aditya Agarwal , Gaurav Singh , Bipasha Sen , Tomás Lozano-Pérez , Leslie Pack Kaelbling

Indoor scene understanding is central to applications such as robot navigation and human companion assistance. Over the last years, data-driven deep neural networks have outperformed many traditional approaches thanks to their…

Computer Vision and Pattern Recognition · Computer Science 2017-07-04 Yinda Zhang , Shuran Song , Ersin Yumer , Manolis Savva , Joon-Young Lee , Hailin Jin , Thomas Funkhouser

Autonomous systems are increasingly deployed in open and dynamic environments -- from city streets to aerial and indoor spaces -- where perception models must remain reliable under sensor noise, environmental variation, and platform shifts.…

Robotics · Computer Science 2026-01-09 Lingdong Kong , Shaoyuan Xie , Zeying Gong , Ye Li , Meng Chu , Ao Liang , Yuhao Dong , Tianshuai Hu , Ronghe Qiu , Rong Li , Hanjiang Hu , Dongyue Lu , Wei Yin , Wenhao Ding , Linfeng Li , Hang Song , Wenwei Zhang , Yuexin Ma , Junwei Liang , Zhedong Zheng , Lai Xing Ng , Benoit R. Cottereau , Wei Tsang Ooi , Ziwei Liu , Zhanpeng Zhang , Weichao Qiu , Wei Zhang , Ji Ao , Jiangpeng Zheng , Siyu Wang , Guang Yang , Zihao Zhang , Yu Zhong , Enzhu Gao , Xinhan Zheng , Xueting Wang , Shouming Li , Yunkai Gao , Siming Lan , Mingfei Han , Xing Hu , Dusan Malic , Christian Fruhwirth-Reisinger , Alexander Prutsch , Wei Lin , Samuel Schulter , Horst Possegger , Linfeng Li , Jian Zhao , Zepeng Yang , Yuhang Song , Bojun Lin , Tianle Zhang , Yuchen Yuan , Chi Zhang , Xuelong Li , Youngseok Kim , Sihwan Hwang , Hyeonjun Jeong , Aodi Wu , Xubo Luo , Erjia Xiao , Lingfeng Zhang , Yingbo Tang , Hao Cheng , Renjing Xu , Wenbo Ding , Lei Zhou , Long Chen , Hangjun Ye , Xiaoshuai Hao , Shuangzhi Li , Junlong Shen , Xingyu Li , Hao Ruan , Jinliang Lin , Zhiming Luo , Yu Zang , Cheng Wang , Hanshi Wang , Xijie Gong , Yixiang Yang , Qianli Ma , Zhipeng Zhang , Wenxiang Shi , Jingmeng Zhou , Weijun Zeng , Kexin Xu , Yuchen Zhang , Haoxiang Fu , Ruibin Hu , Yanbiao Ma , Xiyan Feng , Wenbo Zhang , Lu Zhang , Yunzhi Zhuge , Huchuan Lu , You He , Seungjun Yu , Junsung Park , Youngsun Lim , Hyunjung Shim , Faduo Liang , Zihang Wang , Yiming Peng , Guanyu Zong , Xu Li , Binghao Wang , Hao Wei , Yongxin Ma , Yunke Shi , Shuaipeng Liu , Dong Kong , Yongchun Lin , Huitong Yang , Liang Lei , Haoang Li , Xinliang Zhang , Zhiyong Wang , Xiaofeng Wang , Yuxia Fu , Yadan Luo , Djamahl Etchegaray , Yang Li , Congfei Li , Yuxiang Sun , Wenkai Zhu , Wang Xu , Linru Li , Longjie Liao , Jun Yan , Benwu Wang , Xueliang Ren , Xiaoyu Yue , Jixian Zheng , Jinfeng Wu , Shurui Qin , Wei Cong , Yao He

Many methods in learning from demonstration assume that the demonstrator has knowledge of the full environment. However, in many scenarios, a demonstrator only sees part of the environment and they continuously replan as they gather…

Robotics · Computer Science 2020-05-13 Craig Knuth , Glen Chou , Necmiye Ozay , Dmitry Berenson

Holistic 3D scene understanding involves capturing and parsing unstructured 3D environments. Due to the inherent complexity of the real world, existing models have predominantly been developed and limited to be task-specific. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Sebastian Koch , Johanna Wald , Hidenobu Matsuki , Pedro Hermosilla , Timo Ropinski , Federico Tombari

Navigational signs are common aids for human wayfinding and scene understanding, but are underutilized by robots. We argue that they benefit robot navigation and scene understanding, by directly encoding privileged information on actions,…

Robotics · Computer Science 2025-09-17 Ayush Agrawal , Joel Loo , Nicky Zimmerman , David Hsu