中文
相关论文

相关论文: Is Your Driving World Model an All-Around Player?

200 篇论文

The recent rapid advancement of Text-to-Video (T2V) generation technologies are engaging the trained models with more world model ability, making the existing benchmarks increasingly insufficient to evaluate state-of-the-art T2V models.…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Zeqing Wang , Xinyu Wei , Bairui Li , Zhen Guo , Jinrui Zhang , Hongyang Wei , Keze Wang , Lei Zhang

Current world models lack a unified and controlled setting for systematic evaluation, making it difficult to assess whether they truly capture the underlying rules that govern environment dynamics. In this work, we address this open…

机器学习 · 计算机科学 2025-12-01 Xinyi Li , Zaishuo Xia , Weyl Lu , Chenjie Hao , Yubei Chen

Traditional autonomous driving methods adopt a modular design, decomposing tasks into sub-tasks. In contrast, end-to-end autonomous driving directly outputs actions from raw sensor data, avoiding error accumulation. However, training an…

机器人学 · 计算机科学 2024-11-22 Zeyu Dong , Yimin Zhu , Yansong Li , Kevin Mahon , Yu Sun

Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-vision encoders and Multimodal Large Language Models (MLLMs)…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Xavier Thomas , Youngsun Lim , Ananya Srinivasan , Audrey Zheng , Deepti Ghadiyaram

Realistic and diverse traffic scenarios in large quantities are crucial for the development and validation of autonomous driving systems. However, owing to numerous difficulties in the data collection process and the reliance on intensive…

机器人学 · 计算机科学 2025-10-07 Shuo Sun , Zekai Gu , Tianchen Sun , Jiawei Sun , Chengran Yuan , Yuhang Han , Dongen Li , Marcelo H. Ang

While recent video world models can generate highly realistic videos, their ability to perform semantic reasoning and planning remains unclear and unquantified. We introduce Target-Bench, the first benchmark that enables comprehensive…

World models have become central to autonomous driving, where accurate scene understanding and future prediction are crucial for safe control. Recent work has explored using vision-language models (VLMs) for planning, yet existing…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Zhexiao Xiong , Xin Ye , Burhan Yaman , Sheng Cheng , Yiren Lu , Jingru Luo , Nathan Jacobs , Liu Ren

World models are central to building agents that can reason, plan, and generalize beyond their training data. However, research on world models is currently fragmented, with disparate codebases, data pipelines, and evaluation protocols…

World models, generative AI systems that simulate how environments evolve, are transforming autonomous driving, yet all existing approaches adopt an ego-vehicle perspective, leaving the infrastructure viewpoint unexplored. We argue that…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Siyuan Meng , Chengbo Ai

We introduce LivingWorld, an interactive framework for generating 4D worlds with environmental dynamics from a single image. While recent advances in 3D scene generation enable large-scale environment creation, most approaches focus…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Hyeongju Mun , In-Hwan Jin , Sohyeong Kim , Kyeongbo Kong

Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and traditional simulation suffers from a persistent reality gap. While recent generative…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Kaicong Huang , Talha Azfar , Weisong Shi , Ruimin Ke

World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning. However, current world model research is often dominated by three partially…

人工智能 · 计算机科学 2026-05-27 Sen Cui , Jingheng Ma

With the rapid development of autonomous vehicles, there is an increasing demand for scenario-based testing to simulate diverse driving scenarios. However, as the base of any driving scenarios, road scenarios (e.g., road topology and…

软件工程 · 计算机科学 2024-12-02 Fan Yang , You Lu , Bihuan Chen , Peng Qin , Xin Peng

The generation of realistic and diverse traffic scenarios in simulation is essential for developing and evaluating autonomous driving systems. However, most simulation frameworks rely on rule-based or simplified models for scene generation,…

多智能体系统 · 计算机科学 2025-12-02 Jiaguo Tian , Zhengbang Zhu , Shenyu Zhang , Li Xu , Bo Zheng , Xu Liu , Weiji Peng , Shizeng Yao , Weinan Zhang

World models have become increasingly popular in acting as learned traffic simulators. Recent work has explored replacing traditional traffic simulators with world models for policy training. In this work, we explore the robustness of…

机器人学 · 计算机科学 2025-08-05 Hunter Schofield , Mohammed Elmahgiubi , Kasra Rezaee , Jinjun Shan

Recent advances in text-to-image generation have improved the quality of synthesized images, but evaluations mainly focus on aesthetics or alignment with text prompts. Thus, it remains unclear whether these models can accurately represent a…

Open-world perception aims to develop a model adaptable to novel domains and various sensor configurations and can understand uncommon objects and corner cases. However, current research lacks sufficiently comprehensive open-world 3D…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Zhongyu Xia , Jishuo Li , Zhiwei Lin , Xinhao Wang , Yongtao Wang , Ming-Hsuan Yang

The emergence of Large Language Models (LLMs) has illuminated the potential for a general-purpose user simulator. However, existing benchmarks remain constrained to isolated scenarios, narrow action spaces, or synthetic data, failing to…

Machine learning based autonomous driving systems often face challenges with safety-critical scenarios that are rare in real-world data, hindering their large-scale deployment. While increasing real-world training data coverage could…

机器学习 · 计算机科学 2024-09-13 Yuan Yin , Pegah Khayatan , Éloi Zablocki , Alexandre Boulch , Matthieu Cord

We introduce World Consistency Score (WCS), a novel unified evaluation metric for generative video models that emphasizes internal world consistency of the generated videos. WCS integrates four interpretable sub-components - object…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Akshat Rakheja , Aarsh Ashdhir , Aryan Bhattacharjee , Vanshika Sharma