English

LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation

Computer Vision and Pattern Recognition 2026-05-29 v3 Artificial Intelligence

Abstract

Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a significant challenge. In response to this challenge, we present LoCoT2V-Bench, a benchmark for long video generation (LVG) featuring multi-scene prompts with hierarchical metadata (e.g., character settings and camera behaviors), constructed from collected real-world videos. We further propose LoCoT2V-Eval, a multi-dimensional framework covering perceptual quality, text-video alignment, temporal quality, dynamic quality, and Human Expectation Realization Degree (HERD), with an emphasis on aspects such as fine-grained text-video alignment and temporal character consistency. Experiments on 17 representative LVG models reveal pronounced capability disparities across evaluation dimensions, with strong perceptual quality and background consistency but markedly weaker fine-grained text-video alignment and character consistency. These findings suggest that improving prompt faithfulness and identity preservation remains a key challenge for long-form video generation. Our code and data are released at https://github.com/XqZeppelinhead0702/LoCoT2V-Bench

Keywords

Cite

@article{arxiv.2510.26412,
  title  = {LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation},
  author = {Xiangqing Zheng and Chengyue Wu and Kehai Chen and Min Zhang},
  journal= {arXiv preprint arXiv:2510.26412},
  year   = {2026}
}

Comments

Accepted by ICML 2026 (Regular)

R2 v1 2026-07-01T07:13:41.697Z