English
Related papers

Related papers: AGO: Adaptive Grounding for Open World 3D Occupanc…

200 papers

3D semantic occupancy prediction is one of the crucial tasks of autonomous driving. It enables precise and safe interpretation and navigation in complex environments. Reliable predictions rely on effective sensor fusion, as different…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Tomislav Pavković , Mohammad-Ali Nikouei Mahani , Johannes Niedermayer , Johannes Betz

Occupancy prediction reconstructs 3D structures of surrounding environments. It provides detailed information for autonomous driving planning and navigation. However, most existing methods heavily rely on the LiDAR point clouds to generate…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Chubin Zhang , Juncheng Yan , Yi Wei , Jiaxin Li , Li Liu , Yansong Tang , Yueqi Duan , Jiwen Lu

3D object affordance grounding aims to predict the touchable regions on a 3d object, which is crucial for human-object interaction, human-robot interaction, embodied perception, and robot learning. Recent advances tackle this problem via…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Hanqing Wang , Zhenhao Zhang , Kaiyang Ji , Mingyu Liu , Wenti Yin , Yuchao Chen , Zhirui Liu , Xiangyu Zeng , Tianxiang Gui , Hangxing Zhang

Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Zhihao Yuan , Yibo Peng , Jinke Ren , Yinghong Liao , Yatong Han , Chun-Mei Feng , Hengshuang Zhao , Guanbin Li , Shuguang Cui , Zhen Li

Autonomous driving requires forecasting both geometry and semantics over time to effectively reason about future environment states. Existing vision-based occupancy forecasting methods focus on motion-related categories such as static and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Riya Mohan , Juana Valeria Hurtado , Rohit Mohan , Abhinav Valada

Deformable object manipulation in robotics presents significant challenges due to uncertainties in component properties, diverse configurations, visual interference, and ambiguous prompts. These factors complicate both perception and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Wanjun Jia , Fan Yang , Mengfei Duan , Xianchi Chen , Yinxi Wang , Yiming Jiang , Wenrui Chen , Kailun Yang , Zhiyong Li

Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to leverage commonsense and spatial reasoning for efficient exploration in…

Out-of-distribution (OOD) object detection is a critical task focused on detecting objects that originate from a data distribution different from that of the training data. In this study, we investigate to what extent state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Sadia Ilyas , Ido Freeman , Matthias Rottmann

Occupancy prediction, aiming at predicting the occupancy status within voxelized 3D environment, is quickly gaining momentum within the autonomous driving community. Mainstream occupancy prediction works first discretize the 3D environment…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Jiabao Wang , Zhaojiang Liu , Qiang Meng , Liujiang Yan , Ke Wang , Jie Yang , Wei Liu , Qibin Hou , Ming-Ming Cheng

Online 3D occupancy prediction provides a comprehensive spatial understanding of embodied environments. While the innovative EmbodiedOcc framework utilizes 3D semantic Gaussians for progressive indoor occupancy prediction, it overlooks the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Hao Wang , Xiaobao Wei , Xiaoan Zhang , Jianing Li , Chengyu Bai , Ying Li , Ming Lu , Wenzhao Zheng , Shanghang Zhang

Large foundation models trained on large-scale vision-language data can boost Open-Vocabulary Object Detection (OVD) via synthetic training data, yet the hand-crafted pipelines often introduce bias and overfit to specific prompts. We…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Yang Zhou , Shiyu Zhao , Yuxiao Chen , Zhenting Wang , Can Jin , Dimitris N. Metaxas

The emergence of Vision-Language Models (VLMs) has introduced new paradigms for global image geo-localization through retrieval-augmented generation (RAG) and reasoning-driven inference. However, RAG methods are constrained by retrieval…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Bo Yu , Fengze Yang , Yiming Liu , Chao Wang , Xuewen Luo , Taozhe Li , Ruimin Ke , Xiaofan Zhou , Chenxi Liu

3D semantic occupancy prediction is a pivotal task in the field of autonomous driving. Recent approaches have made great advances in 3D semantic occupancy predictions on a single modality. However, multi-modal semantic occupancy prediction…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Jingyi Pan , Zipeng Wang , Lin Wang

Active perception enables robots to dynamically gather information by adjusting their viewpoints, a crucial capability for interacting with complex, partially observable environments. In this paper, we present AP-VLM, a novel framework that…

Robotics · Computer Science 2025-06-10 Venkatesh Sripada , Samuel Carter , Frank Guerin , Amir Ghalamzan

Future 3D semantic occupancy forecasting and motion planning are central to autonomous driving, as they require models to reason about how surrounding scenes evolve and how the ego vehicle should act. Existing occupancy world models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Cheng Chen , Hao Huang , Saurabh Bagchi

Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-grained language…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Fuhao Li , Huan Jin , Bin Gao , Liaoyuan Fan , Lihui Jiang , Long Zeng

Vision-Language Models (VLMs) exhibit strong visual reasoning capabilities, yet they still struggle with 3D understanding. In particular, VLMs often fail to infer a text-consistent goal 6D pose of a target object in a 3D scene. However, we…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Sangwon Baik , Gunhee Kim , Mingi Choi , Hanbyul Joo

3D visual grounding (3DVG) involves localizing entities in a 3D scene referred to by natural language text. Such models are useful for embodied AI and scene retrieval applications, which involve searching for objects or patterns using…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Austin T. Wang , ZeMing Gong , Angel X. Chang

Perceiving the world and forecasting its future state is a critical task for self-driving. Supervised approaches leverage annotated object labels to learn a model of the world -- traditionally with object detections and trajectory…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Ben Agro , Quinlan Sykora , Sergio Casas , Thomas Gilles , Raquel Urtasun

3D occupancy perception holds a pivotal role in recent vision-centric autonomous driving systems by converting surround-view images into integrated geometric and semantic representations within dense 3D grids. Nevertheless, current models…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Xin Tan , Wenbin Wu , Zhiwei Zhang , Chaojie Fan , Yong Peng , Zhizhong Zhang , Yuan Xie , Lizhuang Ma
‹ Prev 1 4 5 6 7 8 10 Next ›