English

BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios

Multimedia 2026-05-05 v1 Artificial Intelligence Computer Vision and Pattern Recognition

Abstract

The rapid advancement of photorealistic Text-to-Video (T2V) generation brings in an urgent need for up-to-date evaluation methods. Existing benchmarks largely overlooked implausible scenarios and do not measure audio-visual alignment. We introduce BRITE, the first framework that unifies (1) implausible prompting, (2) fine-grained assessment of audio-visual consistency, and (3) QA-based interpretable evaluation into a comprehensive T2V benchmark. Unlike fully automated Multimodal LLM-based pipelines, which are prone to hallucination and prompt ambiguity, BRITE guarantees reliability through a rigorous human-in-the-loop protocol for benchmark creation. Evaluating five state-of-the-art models (Sora 2, Veo 3.1, Runway Gen4.5, Pixverse V5.5, and Qwen3Max), we reveal a critical performance gap: while models excel at static object composition, they exhibit significant degradation in object-action binding and audio-visual synchronization. Our framework offers the community a reliable, interpretable benchmark and evaluation framework that can detect and locate limitations in the next generation of T2V models, especially for off-manifold prompts

Keywords

Cite

@article{arxiv.2605.00873,
  title  = {BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios},
  author = {Advait Tilak and Jiwon Choi and Nazifa Mouli and Wei Le},
  journal= {arXiv preprint arXiv:2605.00873},
  year   = {2026}
}
R2 v1 2026-07-01T12:45:37.168Z