English

HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles

Computer Vision and Pattern Recognition 2026-03-03 v2

Abstract

Controllable driving scene generation is critical for realistic and scalable autonomous driving simulation, yet existing approaches struggle to jointly achieve photorealism and precise control. We introduce HorizonForge, a unified framework that reconstructs scenes as editable Gaussian Splats and Meshes, enabling fine-grained 3D manipulation and language-driven vehicle insertion. Edits are rendered through a noise-aware video diffusion process that enforces spatial and temporal consistency, producing diverse scene variations in a single feed-forward pass without per-trajectory optimization. To standardize evaluation, we further propose HorizonSuite, a comprehensive benchmark spanning ego- and agent-level editing tasks such as trajectory modifications and object manipulation. Extensive experiments show that Gaussian-Mesh representation delivers substantially higher fidelity than alternative 3D representations, and that temporal priors from video diffusion are essential for coherent synthesis. Combining these findings, HorizonForge establishes a simple yet powerful paradigm for photorealistic, controllable driving simulation, achieving an 83.4% user-preference gain and a 25.19% FID improvement over the second best state-of-the-art method. Project page: https://horizonforge.github.io/ .

Keywords

Cite

@article{arxiv.2602.21333,
  title  = {HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles},
  author = {Yifan Wang and Francesco Pittaluga and Zaid Tasneem and Chenyu You and Manmohan Chandraker and Ziyu Jiang},
  journal= {arXiv preprint arXiv:2602.21333},
  year   = {2026}
}

Comments

Accepted by CVPR 2026