English
Related papers

Related papers: GASP: Unifying Geometric and Semantic Self-Supervi…

200 papers

Occupancy prediction infers fine-grained 3D geometry and semantics from camera images of the surrounding environment, making it a critical perception task for autonomous driving. Existing methods either adopt dense grids as scene…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Yunxiao Shi , Yinhao Zhu , Shizhong Han , Jisoo Jeong , Amin Ansari , Hong Cai , Fatih Porikli

Autonomous driving requires forecasting both geometry and semantics over time to effectively reason about future environment states. Existing vision-based occupancy forecasting methods focus on motion-related categories such as static and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Riya Mohan , Juana Valeria Hurtado , Rohit Mohan , Abhinav Valada

Forecasting the scalable future states of surrounding traffic participants in complex traffic scenarios is a critical capability for autonomous vehicles, as it enables safe and feasible decision-making. Recent successes in learning-based…

Robotics · Computer Science 2023-05-08 Haochen Liu , Zhiyu Huang , Chen Lv

Understanding the geometric and semantic structure of environments is essential for embodied navigation and reasoning. Existing semantic mapping methods trade off between explicit geometry and multi-scale semantics, and lack a native…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Sixian Zhang , Yiyao Wang , Xinhang Song , Keming Zhang , Zijian Xu , Shuqiang Jiang

Learning inter-image similarity is crucial for 3D medical images self-supervised pre-training, due to their sharing of numerous same semantic regions. However, the lack of the semantic prior in metrics and the semantic-independent variation…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Yuting He , Guanyu Yang , Rongjun Ge , Yang Chen , Jean-Louis Coatrieux , Boyu Wang , Shuo Li

For autonomous vehicles to proactively plan safe trajectories and make informed decisions, they must be able to predict the future occupancy states of the local environment. However, common issues with occupancy prediction include…

Robotics · Computer Science 2024-04-15 Maneekwan Toyungyernsub , Esen Yel , Jiachen Li , Mykel J. Kochenderfer

Addressing the task of 3D semantic occupancy prediction for autonomous driving, we tackle two key issues in existing 3D Gaussian Splatting (3DGS) methods: (1) unified feature aggregation neglecting semantic correlations among similar…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Ke Song , Yunhe Wu , Chunchit Siu , Huiyuan Xiong

Self-supervised representation learning techniques utilize large datasets without semantic annotations to learn meaningful, universal features that can be conveniently transferred to solve a wide variety of downstream supervised tasks. In…

Computer Vision and Pattern Recognition · Computer Science 2022-10-10 Swetava Ganguli , C. V. Krishnakumar Iyer , Vipul Pandey

Masked signal modeling has greatly advanced self-supervised pre-training for language and 2D images. However, it is still not fully explored in 3D scene understanding. Thus, this paper introduces Masked Shape Prediction (MSP), a new…

Computer Vision and Pattern Recognition · Computer Science 2023-05-10 Li Jiang , Zetong Yang , Shaoshuai Shi , Vladislav Golyanik , Dengxin Dai , Bernt Schiele

In recent years, there has been a rapid development of spatio-temporal prediction techniques in response to the increasing demands of traffic management and travel planning. While advanced end-to-end models have achieved notable success in…

Machine Learning · Computer Science 2023-11-09 Zhonghang Li , Lianghao Xia , Yong Xu , Chao Huang

Being able to perceive the semantics and the spatial structure of the environment is essential for visual navigation of a household robot. However, most existing works only employ visual backbones pre-trained either with independent images…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Yicong Hong , Yang Zhou , Ruiyi Zhang , Franck Dernoncourt , Trung Bui , Stephen Gould , Hao Tan

Autonomous vehicles commonly rely on highly detailed birds-eye-view maps of their environment, which capture both static elements of the scene such as road layout as well as dynamic elements such as other cars and pedestrians. Generating…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Thomas Roddick , Roberto Cipolla

Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relies on large amounts of annotated data. Recent self-supervised approaches, such as Masked…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Zhuohao Chen , Zeng Li , Yifei Zhang , Chang Liu , Yu Zhou

Accurate, real-time 3D reconstruction of human heads from monocular images and videos underlies numerous visual applications. As 3D ground truth data is hard to come by at scale, previous methods have sought to learn from abundant 2D videos…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Liam Schoneveld , Zhe Chen , Davide Davoli , Jiapeng Tang , Saimon Terazawa , Ko Nishino , Matthias Nießner

Fast, collision-free motion through unknown environments remains a challenging problem for robotic systems. In these situations, the robot's ability to reason about its future motion is often severely limited by sensor field of view (FOV).…

Machine Learning · Computer Science 2018-03-07 Kapil Katyal , Katie Popek , Chris Paxton , Joseph Moore , Kevin Wolfe , Philippe Burlina , Gregory D. Hager

Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases, while integrating specialized…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Chun-Hsiao Yeh , Shengyi Qian , Manchen Wang , Yi Ma , Joseph Tighe , Fanyi Xiao

The 3D occupancy prediction task has witnessed remarkable progress in recent years, playing a crucial role in vision-based autonomous driving systems. While traditional methods are limited to fixed semantic categories, recent approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Chi Yan , Dan Xu

Self-supervision can dramatically cut back the amount of manually-labelled data required to train deep neural networks. While self-supervision has usually been considered for tasks such as image classification, in this paper we aim at…

Computer Vision and Pattern Recognition · Computer Science 2018-04-06 David Novotny , Samuel Albanie , Diane Larlus , Andrea Vedaldi

Witnessing the impressive achievements of pre-training techniques on large-scale data in the field of computer vision and natural language processing, we wonder whether this idea could be adapted in a grab-and-go spirit, and mitigate the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Penghao Wu , Li Chen , Hongyang Li , Xiaosong Jia , Junchi Yan , Yu Qiao

We present a general framework to autonomously achieve a task, where autonomy is acquired by learning sensorimotor patterns of a robot, while it is interacting with its environment. To accomplish the task, using the learned sensorimotor…

Robotics · Computer Science 2016-01-06 Ali Ghadirzadeh , Judith Bütepage , Danica Kragic , Mårten Björkman