中文
相关论文

相关论文: GRADEO: Towards Human-Like Evaluation for Text-to-…

200 篇论文

Multimodal generative models have made significant strides in image editing, demonstrating impressive performance on a variety of static tasks. However, their proficiency typically does not extend to complex scenarios requiring dynamic…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Zhiqiang Sheng , Xumeng Han , Zhiwei Zhang , Zenghui Xiong , Yifan Ding , Aoxiang Ping , Xiang Li , Tong Guo , Yao Mao

The growing demand for high-fidelity video generation from textual descriptions has catalyzed significant research in this field. In this work, we introduce MagicVideo-V2 that integrates the text-to-image model, video motion generator,…

计算机视觉与模式识别 · 计算机科学 2024-01-10 Weimin Wang , Jiawei Liu , Zhijie Lin , Jiangqiao Yan , Shuo Chen , Chetwin Low , Tuyen Hoang , Jie Wu , Jun Hao Liew , Hanshu Yan , Daquan Zhou , Jiashi Feng

While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While previous work has evaluated T2I alignment by proposing metrics, benchmarks, and templates for…

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Weixia Zhang , Chengguang Zhu , Jingnan Gao , Yichao Yan , Guangtao Zhai , Xiaokang Yang

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benchmarks mostly resemble static image benchmarks and are…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Mu Cai , Reuben Tan , Jianrui Zhang , Bocheng Zou , Kai Zhang , Feng Yao , Fangrui Zhu , Jing Gu , Yiwu Zhong , Yuzhang Shang , Yao Dou , Jaden Park , Jianfeng Gao , Yong Jae Lee , Jianwei Yang

Contemporary Text-to-Image (T2I) models frequently depend on qualitative human evaluations to assess the consistency between synthesized images and the text prompts. There is a demand for quantitative and automatic evaluation tools, given…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Ziyuan Qin , Dongjie Cheng , Haoyu Wang , Huahui Yi , Yuting Shao , Zhiyuan Fan , Kang Li , Qicheng Lao

Video color grading is a critical post-production process that transforms flat, log-encoded raw footage into emotionally resonant cinematic visuals. Existing automated methods act as static, black-box executors that directly output edited…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Yuchen Guo , Junli Gong , Hongmin Cai , Yiu-ming Cheung , Weifeng Su

The current state-of-the-art video generative models can produce commercial-grade videos with highly realistic details. However, they still struggle to coherently present multiple sequential events in the stories specified by the prompts,…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Yiping Wang , Xuehai He , Kuan Wang , Luyao Ma , Jianwei Yang , Shuohang Wang , Simon Shaolei Du , Yelong Shen

Large language models show improved downstream task performance when prompted to generate step-by-step reasoning to justify their final answers. These reasoning steps greatly improve model interpretability and verification, but objectively…

Despite recent advances in text-to-3D generative methods, there is a notable absence of reliable evaluation metrics. Existing metrics usually focus on a single criterion each, such as how well the asset aligned with the input text. These…

计算机视觉与模式识别 · 计算机科学 2024-01-11 Tong Wu , Guandao Yang , Zhibing Li , Kai Zhang , Ziwei Liu , Leonidas Guibas , Dahua Lin , Gordon Wetzstein

Text-driven video editing enables users to modify video content only using text queries. While existing methods can modify video content if explicit descriptions of editing targets with precise spatial locations and temporal boundaries are…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Yiqing Shen , Chenjia Li , Mathias Unberath

Advancements in language foundation models have primarily fueled the recent surge in artificial intelligence. In contrast, generative learning of non-textual modalities, especially videos, significantly trails behind language modeling. This…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Lijun Yu

Despite significant breakthroughs in video analysis driven by the rapid development of large multimodal models (LMMs), there remains a lack of a versatile evaluation benchmark to comprehensively assess these models' performance in video…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Yunxin Li , Xinyu Chen , Baotian Hu , Longyue Wang , Haoyuan Shi , Min Zhang

Video generation models have achieved remarkable progress in text-to-video tasks. These models are typically trained on text-video pairs with highly detailed and carefully crafted descriptions, while real-world user inputs during inference…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Jiale Cheng , Ruiliang Lyu , Xiaotao Gu , Xiao Liu , Jiazheng Xu , Yida Lu , Jiayan Teng , Zhuoyi Yang , Yuxiao Dong , Jie Tang , Hongning Wang , Minlie Huang

Retrieval-Augmented Generation (RAG) systems are widely adopted in knowledge-intensive NLP tasks, but current evaluations often overlook the structural complexity and multi-step reasoning required in real-world scenarios. These benchmarks…

计算与语言 · 计算机科学 2025-12-16 Jeongsoo Lee , Daeyong Kwon , Kyohoon Jin

Video matting has traditionally been limited by the lack of high-quality ground-truth data. Most existing video matting datasets provide only human-annotated imperfect alpha and foreground annotations, which must be composited to background…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Yongtao Ge , Kangyang Xie , Guangkai Xu , Mingyu Liu , Li Ke , Longtao Huang , Hui Xue , Hao Chen , Chunhua Shen

Automatically generating scripts (i.e. sequences of key steps described in text) from video demonstrations and reasoning about the subsequent steps are crucial to the modern AI virtual assistants to guide humans to complete everyday tasks,…

计算与语言 · 计算机科学 2024-01-22 Jingyuan Qi , Minqian Liu , Ying Shen , Zhiyang Xu , Lifu Huang

We show that video generation models could reason now. Testing on tasks such as chess, maze, Sudoku, mental rotation, and Raven's Matrices, leading models such as Sora-2 achieve sixty percent success rates. We establish a robust…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Hokin Deng

While large-scale datasets have driven significant progress in Text-to-Video (T2V) generative models, these models remain highly sensitive to input prompts, demonstrating that prompt design is critical to generation quality. Current methods…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Zillur Rahman , Alex Sheng , Cristian Meo