English
Related papers

Related papers: Infrastructure-Centric World Models: Bridging Temp…

200 papers

Perceiving the world and forecasting its future state is a critical task for self-driving. Supervised approaches leverage annotated object labels to learn a model of the world -- traditionally with object detections and trajectory…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Ben Agro , Quinlan Sykora , Sergio Casas , Thomas Gilles , Raquel Urtasun

This paper introduces a multi-agent framework for comprehensive highway scene understanding, designed around a mixture-of-experts strategy. In this framework, a large generic vision-language model (VLM), such as GPT-4o, is contextualized…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Yunxiang Yang , Ningning Xu , Jidong J. Yang

Navigation is a fundamental capability for mobile robots. While the current trend is to use learning-based approaches to replace traditional geometry-based methods, existing end-to-end learning-based policies often struggle with 3D spatial…

Robotics · Computer Science 2026-01-21 Wangtian Shen , Ziyang Meng , Jinming Ma , Mingliang Zhou , Diyun Xiang

In autonomous driving, transparency in the decision-making of perception models is critical, as even a single misperception can be catastrophic. Yet with multi-sensor inputs, it is difficult to determine how each modality contributes to a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Jaehyun Park , Konyul Park , Daehun Kim , Junseo Park , Jun Won Choi

We present INDOOR-LIDAR, a comprehensive hybrid dataset of indoor 3D LiDAR point clouds designed to advance research in robot perception. Existing indoor LiDAR datasets often suffer from limited scale, inconsistent annotation formats, and…

Robotics · Computer Science 2025-12-16 Haichuan Li , Changda Tian , Panos Trahanias , Tomi Westerlund

Today's autonomous vehicles rely extensively on high-definition 3D maps to navigate the environment. While this approach works well when these maps are completely up-to-date, safe autonomous vehicles must be able to corroborate the map's…

Computer Vision and Pattern Recognition · Computer Science 2016-12-09 Ari Seff , Jianxiong Xiao

Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive world modeling enables VLA to internalize physical dynamics…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Baolu Li , Jingyu Qian , Rui Guo , Yilun Chen , Hanpeng Liu , Yuan Lin , Junhong Zhou , Ruixin Liu , Willow Yang , Yutong Zheng , Zhenli Zhang , Tenglong , Gu , Zhuangzhuang Ding , Pengkun Zheng , Yu Zhang , Xianming Liu

The vision-language-action (VLA) paradigm has enabled powerful robotic control by leveraging vision-language models, but its reliance on large-scale, high-quality robot data limits its generalization. Generative world models offer a…

Robotics · Computer Science 2026-01-27 Weishi Mi , Yong Bao , Xiaowei Chi , Xiaozhu Ju , Zhiyuan Qin , Kuangzhi Ge , Kai Tang , Peidong Jia , Shanghang Zhang , Jian Tang

For autonomous driving, traversability analysis is one of the most basic and essential tasks. In this paper, we propose a novel LiDAR-based terrain modeling approach, which could output stable, complete and accurate terrain models and…

Robotics · Computer Science 2023-07-06 Hanzhang Xue , Hao Fu , Liang Xiao , Yiming Fan , Dawei Zhao , Bin Dai

Automation of complex traffic scenarios is expected to rely on input from a roadside infrastructure to complement the vehicles' environment perception. We here explore design requirements for a prototypical setup of virtual vision or RADAR…

Signal Processing · Electrical Eng. & Systems 2019-02-26 Florian Geissler , Sören Kohnert , Reinhard Stolle

Fusing different sensor modalities can be a difficult task, particularly if they are asynchronous. Asynchronisation may arise due to long processing times or improper synchronisation during calibration, and there must exist a way to still…

Robotics · Computer Science 2024-10-02 Seamie Hayes , Sushil Sharma , Ciarán Eising

Detection and segmentation of moving obstacles, along with prediction of the future occupancy states of the local environment, are essential for autonomous vehicles to proactively make safe and informed decisions. In this paper, we propose…

Robotics · Computer Science 2022-09-28 Maneekwan Toyungyernsub , Esen Yel , Jiachen Li , Mykel J. Kochenderfer

Embodied navigation in open, dynamic environments demands accurate foresight of how the world will evolve and how actions will unfold over time. We propose AstraNav-World, an end-to-end world model that jointly reasons about future visual…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Jintao Chen , Junjun Hu , Haochen Bai , Minghua Luo , Xinda Xue , Botao Ren , Chengyu Bai , Shichao Xie , Ziyi Chen , Fei Liu , Zedong Chu , Xiaolong Wu , Mu Xu , Shanghang Zhang

Predicting future trajectories of traffic agents in highly interactive environments is an essential and challenging problem for the safe operation of autonomous driving systems. On the basis of the fact that self-driving vehicles are…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Chiho Choi , Joon Hee Choi , Jiachen Li , Srikanth Malla

Predicting future trajectories of traffic agents in highly interactive environments is an essential and challenging problem for the safe operation of autonomous driving systems. On the basis of the fact that self-driving vehicles are…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Chiho Choi , Joon Hee Choi , Srikanth Malla , Jiachen Li

Ophthalmic decision-making depends on subtle lesion-scale cues interpreted across multimodal imaging and over time, yet most medical foundation models remain static and degrade under modality and acquisition shifts. Here we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ziyu Gao , Xinyuan Wu , Xiaolan Chen , Zhuoran Liu , Ruoyu Chen , Bowen Liu , Bingjie Yan , Zhenhan Wang , Kai Jin , Jiancheng Yang , Yih Chung Tham , Mingguang He , Danli Shi

The Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Mingliang Zhai , Cheng Li , Zengyuan Guo , Ningrui Yang , Xiameng Qin , Sanyuan Zhao , Junyu Han , Ji Tao , Yuwei Wu , Yunde Jia

World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Xinyu Zhang , Zhengtong Xu , Yutian Tao , Yeping Wang , Yu She , Abdeslam Boularias

This paper discusses the limitations of existing microscopic traffic models in accounting for the potential impacts of on-ramp vehicles on the car-following behavior of main-lane vehicles on highways. We first surveyed U.S. on-ramps to…

Systems and Control · Electrical Eng. & Systems 2023-05-23 Dustin Holley , Jovin D'sa , Hossein Nourkhiz Mahjoub , Gibran Ali , Behdad Chalaki , Ehsan Moradi-Pari

Spatial synchronization in roadside scenarios is essential for integrating data from multiple sensors at different locations. Current methods using cascading spatial transformation (CST) often lead to cumulative errors in large-scale…

Signal Processing · Electrical Eng. & Systems 2023-11-09 Yong Li , Zhiguo Zhao , Yunli Chen , Rui Tian
‹ Prev 1 8 9 10 Next ›