End-to-end Conditional Diffusion for Realistic and Controllable Visual Traffic Scenario Generation
Abstract
Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions. Existing learning-based methods often struggle to balance controllability and realism, offering either limited fine-grained control over traffic behavior or controllable scenarios at the expense of behavioral plausibility. This paper presents E2E-CDiff, an end-to-end conditional diffusion framework for controllable and realistic scenario generation. Conditioned on front-view visual observations, E2E-CDiff jointly denoises future motion states and executable low-level controls for route-interacting background vehicles. This unified state-action generation mitigates the planning-control mismatch in conventional two-stage trajectory-then-controller pipelines. Differentiable guidance further regulates speed, enforces drivable-area compliance, and supports collision-avoidance or collision-seeking behaviors, enabling both naturalistic and safety-critical scenario generation. Experiments on Bench2Drive show that E2E-CDiff achieves a favorable controllability-realism trade-off compared with representative reinforcement- and imitation-learning baselines, while its collision-guided variant induces challenging interactions across multiple autonomous driving systems. E2E-CDiff also performs competitively as a learning-based ego planner, demonstrating the generality of end-to-end state-action diffusion.
Cite
@article{arxiv.2607.18637,
title = {End-to-end Conditional Diffusion for Realistic and Controllable Visual Traffic Scenario Generation},
author = {Jingzheng Li and Yufei Ge and Zhijun Chen and Qianren Mao and Zizhe Wang and Binhang Qi and Bing Li and Keyu Chen and Baochang Zhang and Xianglong Liu and Philip S Yu},
journal= {arXiv preprint arXiv:2607.18637},
year = {2026}
}