English

OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving

Computer Vision and Pattern Recognition 2025-12-17 v1

Abstract

Autonomous driving has seen remarkable advancements, largely driven by extensive real-world data collection. However, acquiring diverse and corner-case data remains costly and inefficient. Generative models have emerged as a promising solution by synthesizing realistic sensor data. However, existing approaches primarily focus on single-modality generation, leading to inefficiencies and misalignment in multimodal sensor data. To address these challenges, we propose OminiGen, which generates aligned multimodal sensor data in a unified framework. Our approach leverages a shared Bird\u2019s Eye View (BEV) space to unify multimodal features and designs a novel generalizable multimodal reconstruction method, UAE, to jointly decode LiDAR and multi-view camera data. UAE achieves multimodal sensor decoding through volume rendering, enabling accurate and flexible reconstruction. Furthermore, we incorporate a Diffusion Transformer (DiT) with a ControlNet branch to enable controllable multimodal sensor generation. Our comprehensive experiments demonstrate that OminiGen achieves desired performances in unified multimodal sensor data generation with multimodal consistency and flexible sensor adjustments.

Keywords

Cite

@article{arxiv.2512.14225,
  title  = {OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving},
  author = {Tao Tang and Enhui Ma and xia zhou and Letian Wang and Tianyi Yan and Xueyang Zhang and Kun Zhan and Peng Jia and XianPeng Lang and Jia-Wang Bian and Kaicheng Yu and Xiaodan Liang},
  journal= {arXiv preprint arXiv:2512.14225},
  year   = {2025}
}

Comments

ACM MM 2025

R2 v1 2026-07-01T08:27:03.800Z