English
Related papers

Related papers: T2VTextBench: A Human Evaluation Benchmark for Tex…

200 papers

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models to learn the…

Sound · Computer Science 2024-03-14 Shentong Mo , Jing Shi , Yapeng Tian

We propose a novel text-to-video (T2V) generation benchmark, ChronoMagic-Bench, to evaluate the temporal and metamorphic capabilities of the T2V models (e.g. Sora and Lumiere) in time-lapse video generation. In contrast to existing…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Shenghai Yuan , Jinfa Huang , Yongqi Xu , Yaoyang Liu , Shaofeng Zhang , Yujun Shi , Ruijie Zhu , Xinhua Cheng , Jiebo Luo , Li Yuan

Text-to-audio-video (T2AV) generation is central to applications such as filmmaking and world modeling. However, current models often fail to produce physically plausible sounds. Previous benchmarks primarily focus on audio-video temporal…

The burgeoning field of Artificial Intelligence Generated Content (AIGC) is witnessing rapid advancements, particularly in video generation. This paper introduces AIGCBench, a pioneering comprehensive and scalable benchmark designed to…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Fanda Fan , Chunjie Luo , Wanling Gao , Jianfeng Zhan

Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic…

Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks-largely designed for realistic videos-struggle to evaluate animation-style generation with its…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Leyi Wu , Pengjun Fang , Kai Sun , Yazhou Xing , Yinwei Wu , Songsong Wang , Ziqi Huang , Dan Zhou , Yingqing He , Ying-Cong Chen , Qifeng Chen

Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. This paper presents…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jing Gu , Xian Liu , Yu Zeng , Ashwin Nagarajan , Fangrui Zhu , Daniel Hong , Yue Fan , Qianqi Yan , Kaiwen Zhou , Ming-Yu Liu , Xin Eric Wang

The evolution of video generation toward complex, multi-shot narratives has exposed a critical deficit in current evaluation methods. Existing benchmarks remain anchored to single-shot paradigms, lacking the comprehensive story assets and…

Multimedia · Computer Science 2026-03-02 Haoyuan Shi , Yunxin Li , Nanhao Deng , Zhenran Xu , Xinyu Chen , Longyue Wang , Baotian Hu , Min Zhang

This paper introduces ModelScopeT2V, a text-to-video synthesis model that evolves from a text-to-image synthesis model (i.e., Stable Diffusion). ModelScopeT2V incorporates spatio-temporal blocks to ensure consistent frame generation and…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Jiuniu Wang , Hangjie Yuan , Dayou Chen , Yingya Zhang , Xiang Wang , Shiwei Zhang

Recent advances in video generation have been remarkable, enabling models to produce visually compelling videos with synchronized audio. While existing video generation benchmarks provide comprehensive metrics for visual quality, they lack…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Daili Hua , Xizhi Wang , Bohan Zeng , Xinyi Huang , Hao Liang , Junbo Niu , Xinlong Chen , Quanqing Xu , Wentao Zhang

Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order. Despite the emerging advancements in interleaved generation, the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Minqian Liu , Zhiyang Xu , Zihao Lin , Trevor Ashby , Joy Rimchala , Jiaxin Zhang , Lifu Huang

Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactions. Yet evaluation for AV generation remains at an early…

Artificial Intelligence · Computer Science 2026-05-26 Jialiang Yang , Bin Xia , Ruihang Chu , Dingdong Wang , Wanke Xia , Zhun Mou , Tianyang Zhong , Yiting Zhao , Wenming Yang

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Yuhang Yang , Ke Fan , Shangkun Sun , Hongxiang Li , Ailing Zeng , FeiLin Han , Wei Zhai , Wei Liu , Yang Cao , Zheng-Jun Zha

Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text-video alignment, or physical…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Xianjing Han , Bin Zhu , Shiqi Hu , Franklin Mingzhe Li , Patrick Carrington , Roger Zimmermann , Jingjing Chen

Video generation has advanced significantly, evolving from producing unrealistic outputs to generating videos that appear visually convincing and temporally coherent. To evaluate these video generative models, benchmarks such as VBench have…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Dian Zheng , Ziqi Huang , Hongbo Liu , Kai Zou , Yinan He , Fan Zhang , Lulu Gu , Yuanhan Zhang , Jingwen He , Wei-Shi Zheng , Yu Qiao , Ziwei Liu

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Zhe Cao , Tao Wang , Jiaming Wang , Yanghai Wang , Yuanxing Zhang , Jialu Chen , Miao Deng , Jiahao Wang , Yubin Guo , Chenxi Liao , Yize Zhang , Zhaoxiang Zhang , Jiaheng Liu

The Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both audio and video elements. To support research in this field,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Yuxin Mao , Xuyang Shen , Jing Zhang , Zhen Qin , Jinxing Zhou , Mochu Xiang , Yiran Zhong , Yuchao Dai

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Jiasong Feng , Ao Ma , Jing Wang , Ke Cao , Zhanjie Zhang

Automated data visualization plays a crucial role in simplifying data interpretation, enhancing decision-making, and improving efficiency. While large language models (LLMs) have shown promise in generating visualizations from natural…

Computation and Language · Computer Science 2025-07-29 Mizanur Rahman , Md Tahmid Rahman Laskar , Shafiq Joty , Enamul Hoque

Layout-guided text-to-image models offer greater control over the generation process by explicitly conditioning image synthesis on the spatial arrangement of elements. As a result, their adoption has increased in many computer vision…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Elena Izzo , Luca Parolari , Davide Vezzaro , Lamberto Ballan