English

ARCON: Advancing Auto-Regressive Continuation for Driving Videos

Computer Vision and Pattern Recognition 2025-02-27 v3

Abstract

Recent advancements in auto-regressive large language models (LLMs) have led to their application in video generation. This paper explores the use of Large Vision Models (LVMs) for video continuation, a task essential for building world models and predicting future frames. We introduce ARCON, a scheme that alternates between generating semantic and RGB tokens, allowing the LVM to explicitly learn high-level structural video information. We find high consistency in the RGB images and semantic maps generated without special design. Moreover, we employ an optical flow-based texture stitching method to enhance visual quality. Experiments in autonomous driving scenarios show that our model can consistently generate long videos.

Keywords

Cite

@article{arxiv.2412.03758,
  title  = {ARCON: Advancing Auto-Regressive Continuation for Driving Videos},
  author = {Ruibo Ming and Jingwei Wu and Zhewei Huang and Zhuoxuan Ju and Jianming HU and Lihui Peng and Shuchang Zhou},
  journal= {arXiv preprint arXiv:2412.03758},
  year   = {2025}
}
R2 v1 2026-06-28T20:23:36.576Z