English
Related papers

Related papers: T2VWorldBench: A Benchmark for Evaluating World Kn…

200 papers

Reasoning is a fundamental capability often required in real-world text-to-image (T2I) generation, e.g., generating ``a bitten apple that has been left in the air for more than a week`` necessitates understanding temporal decay and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Kaijie Chen , Zihao Lin , Zhiyang Xu , Ying Shen , Yuguang Yao , Joy Rimchala , Jiaxin Zhang , Lifu Huang

Current multimodal models aim to transcend the limitations of single-modality representations by unifying understanding and generation, often using text-to-image (T2I) tasks to calibrate semantic consistency. However, their reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Juanxi Tian , Siyuan Li , Conghui He , Lijun Wu , Cheng Tan

Video generation models are rapidly advancing, but can still struggle with complex video outputs that require significant semantic branching or repeated high-level reasoning about what should happen next. In this paper, we introduce a new…

Large language models have demonstrated impressive performance when integrated with vision models even enabling video understanding. However, evaluating video models presents its own unique challenges, for which several benchmarks have been…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Daniel Cores , Michael Dorkenwald , Manuel Mucientes , Cees G. M. Snoek , Yuki M. Asano

Video-to-Audio (V2A) generation is essential for immersive multimedia experiences, yet its evaluation remains underexplored. Existing benchmarks typically assess diverse audio types under a unified protocol, overlooking the fine-grained…

Sound · Computer Science 2026-04-14 Qian Zhang , Yuqin Cao , Yixuan Gao , Xiongkuo Min

Existing evaluation frameworks for Multimodal Large Language Models (MLLMs) primarily focus on image reasoning or general video understanding tasks, largely overlooking the significant role of image context in video comprehension. To bridge…

With the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However, most benchmarks predominantly assess…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Kunchang Li , Yali Wang , Yinan He , Yizhuo Li , Yi Wang , Yi Liu , Zun Wang , Jilan Xu , Guo Chen , Ping Luo , Limin Wang , Yu Qiao

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Jianhui Wei , Xiaotian Zhang , Yichen Li , Yuan Wang , Yan Zhang , Ziyi Chen , Zhihang Tang , Wei Xu , Zuozhu Liu

Top-down images play an important role in safety-critical settings such as autonomous navigation and aerial surveillance, where they provide holistic spatial information that front-view images cannot capture. Despite this, Vision Language…

Machine Learning · Computer Science 2025-10-02 Kaiyuan Hou , Minghui Zhao , Lilin Xu , Yuang Fan , Xiaofan Jiang

With the rapid advancement of video understanding, existing benchmarks are becoming increasingly saturated, exposing a critical discrepancy between inflated leaderboard scores and real-world model capabilities. To address this widening gap,…

Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Weihan Wang , Zehai He , Wenyi Hong , Yean Cheng , Xiaohan Zhang , Ji Qi , Xiaotao Gu , Shiyu Huang , Bin Xu , Yuxiao Dong , Ming Ding , Jie Tang

The Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both audio and video elements. To support research in this field,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Yuxin Mao , Xuyang Shen , Jing Zhang , Zhen Qin , Jinxing Zhou , Mochu Xiang , Yiran Zhong , Yuchao Dai

Text-to-Visualization (Text2VIS) enables users to create visualizations from natural language queries, making data insights more accessible. However, Text2VIS faces challenges in interpreting ambiguous queries, as users often express their…

Computation and Language · Computer Science 2026-01-06 Tianqi Luo , Chuhan Huang , Leixian Shen , Boyan Li , Shuyu Shen , Wei Zeng , Nan Tang , Yuyu Luo

Recent text-to-video (T2V) technology advancements, as demonstrated by models such as Gen2, Pika, and Sora, have significantly broadened its applicability and popularity. Despite these strides, evaluating these models poses substantial…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Tianle Zhang , Langtian Ma , Yuchen Yan , Yuchen Zhang , Kai Wang , Yue Yang , Ziyao Guo , Wenqi Shao , Yang You , Yu Qiao , Ping Luo , Kaipeng Zhang

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Yuhang Yang , Ke Fan , Shangkun Sun , Hongxiang Li , Ailing Zeng , FeiLin Han , Wei Zhai , Wei Liu , Yang Cao , Zheng-Jun Zha

Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. While recent Large Multimodal Models (LMMs) have shown…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Andong Deng , Dawei Du , Zhenfang Chen , Wen Zhong , Fan Chen , Guang Chen , Chia-Wen Kuo , Longyin Wen , Chen Chen , Sijie Zhu

While recent video world models can generate highly realistic videos, their ability to perform semantic reasoning and planning remains unclear and unquantified. We introduce Target-Bench, the first benchmark that enables comprehensive…

Text-to-video (T2V) generation has been recently enabled by transformer-based diffusion models, but current T2V models lack capabilities in adhering to the real-world common knowledge and physical rules, due to their limited understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Qiyao Xue , Xiangyu Yin , Boyuan Yang , Wei Gao

In this paper, we present an empirical study introducing a nuanced evaluation framework for text-to-image (T2I) generative models, applied to human image synthesis. Our framework categorizes evaluations into two distinct groups: first,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Muxi Chen , Yi Liu , Jian Yi , Changran Xu , Qiuxia Lai , Hongliang Wang , Tsung-Yi Ho , Qiang Xu

Synthetic videos nowadays is widely used to complement data scarcity and diversity of real-world videos. Current synthetic datasets primarily replicate real-world scenarios, leaving impossible, counterfactual and anti-reality video concepts…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Zechen Bai , Hai Ci , Mike Zheng Shou
‹ Prev 1 4 5 6 7 8 10 Next ›