English
Related papers

Related papers: DirectorBench: Diagnosing Long-Form Video Generati…

200 papers

Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Keming Wu , Yijing Cui , Wenhan Xue , Qijie Wang , Xuan Luo , Zhiyuan Feng , Zuhao Yang , Sudong Wang , Sicong Jiang , Haowei Zhu , Zihan Wang , Ping Nie , Wenhu Chen , Bin Wang

Generating long videos that can show complex stories, like movie scenes from scripts, has great promise and offers much more than short clips. However, current methods that use autoregression with diffusion models often struggle because…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Guangcong Zheng , Jianlong Yuan , Bo Wang , Haoyang Huang , Guoqing Ma , Nan Duan

Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Wenqi Ouyang , Zeqi Xiao , Danni Yang , Yifan Zhou , Shuai Yang , Lei Yang , Jianlou Si , Xingang Pan

With the rapid development of foundation video generation technologies, long video generation models have exhibited promising research potential thanks to expanded content creation space. Recent studies reveal that the goal of long video…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 X. Feng , H. Yu , M. Wu , S. Hu , J. Chen , C. Zhu , J. Wu , X. Chu , K. Huang

Current benchmarks like Needle-in-a-Haystack (NIAH), Ruler, and Needlebench focus on models' ability to understand long-context input sequences but fail to capture a critical dimension: the generation of high-quality long-form text.…

Computation and Language · Computer Science 2025-01-24 Yuhao Wu , Ming Shan Hee , Zhiqing Hu , Roy Ka-Wei Lee

The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interactive videos, such as movies and recordings. To fill this gap,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Xiaodong Wang , Langling Huang , Zhirong Wu , Xu Zhao , Teng Xu , Xuhong Xia , Peixi Peng

As Vision-Language Models (VLMs) increasingly gain traction in medical applications, clinicians are progressively expecting AI systems not only to generate textual diagnoses but also to produce corresponding medical images that integrate…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Junjie Yang , Yuhao Yan , Gang Wu , Yuxuan Wang , Ruoyu Liang , Xinjie Jiang , Xiang Wan , Fenglei Fan , Yongquan Zhang , Feiwei Qin , Changmiao Wang

Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute actions to the correct…

Computation and Language · Computer Science 2026-04-14 Yunzhe Wang , Runhui Xu , Kexin Zheng , Tianyi Zhang , Jayavibhav Niranjan Kogundi , Soham Hans , Volkan Ustun

In this paper, we introduce DirectorLLM, a novel video generation model that employs a large language model (LLM) to orchestrate human poses within videos. As foundational text-to-video models rapidly evolve, the demand for high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Kunpeng Song , Tingbo Hou , Zecheng He , Haoyu Ma , Jialiang Wang , Animesh Sinha , Sam Tsai , Yaqiao Luo , Xiaoliang Dai , Li Chen , Xide Xia , Peizhao Zhang , Peter Vajda , Ahmed Elgammal , Felix Juefei-Xu

In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed across the open web; models therefore need to jointly perform cross-frame clue extraction, iterative…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Chengwen Liu , Xiaomin Yu , Zhuoyue Chang , Zhe Huang , Shuo Zhang , Heng Lian , Jisheng Dang , Rui Xu , Sen Hu , Jianheng Hou , Chengwei Qin , Xiaobin Hu , Kunyi Wang , Zhi Yang , Hao Peng , Hong Peng , Ronghao Chen , Huacan Wang

How do two individuals differ when performing the same action? In this work, we introduce Video Action Differencing (VidDiff), the novel task of identifying subtle differences between videos of the same action, which has many applications,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 James Burgess , Xiaohan Wang , Yuhui Zhang , Anita Rau , Alejandro Lozano , Lisa Dunlap , Trevor Darrell , Serena Yeung-Levy

Data analytics is essential for extracting valuable insights from data that can assist organizations in making effective decisions. We introduce InsightBench, a benchmark dataset with three key features. First, it consists of 100 datasets…

Existing long-form video generation frameworks lack automated planning, requiring manual input for storylines, scenes, cinematography, and character interactions, resulting in high costs and inefficiencies. To address these challenges, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Weijia Wu , Zeyu Zhu , Mike Zheng Shou

Over the years, performance evaluation has become essential in computer vision, enabling tangible progress in many sub-fields. While talking-head video generation has become an emerging research topic, existing evaluations on this topic…

Computer Vision and Pattern Recognition · Computer Science 2020-05-08 Lele Chen , Guofeng Cui , Ziyi Kou , Haitian Zheng , Chenliang Xu

Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users. While emerging benchmarks address proactivity, they are largely confined to alert…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Dongchuan Ran , Linyu Ou , Xueheng Li , Wenwen Tong , Chenxu Guo , Hewei Guo , Kaibing Wang , Lewei Lu

The recent renaissance in generative models, driven primarily by the advent of diffusion models and iterative improvement in GAN methods, has enabled many creative applications. However, each advancement is also accompanied by a rise in the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-28 Sanjay Saha , Rashindrie Perera , Sachith Seneviratne , Tamasha Malepathirana , Sanka Rasnayaka , Deshani Geethika , Terence Sim , Saman Halgamuge

We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step…

Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-30 Dapeng Wu , Shun Lei , Wei Tan , Guangzheng Li , Yunzhe Wang , Huaicheng Zhang , Lishi Zuo , Zhiyong Wu

The recent years have witnessed great advances in video generation. However, the development of automatic video metrics is lagging significantly behind. None of the existing metric is able to provide reliable scores over generated videos.…

Despite rapid advancements in video generation models, generating coherent storytelling videos that span multiple scenes and characters remains challenging. Current methods often rigidly convert pre-generated keyframes into fixed-length…

Multiagent Systems · Computer Science 2025-10-03 Haoyuan Shi , Yunxin Li , Xinyu Chen , Longyue Wang , Baotian Hu , Min Zhang