English
Related papers

Related papers: MMVU: Measuring Expert-Level Multi-Discipline Vide…

200 papers

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the absence of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Yuansen Liu , Haiming Tang , Jinlong Peng , Jiangning Zhang , Xiaozhong Ji , Qingdong He , Wenbin Wu , Donghao Luo , Zhenye Gan , Junwei Zhu , Yunhang Shen , Chaoyou Fu , Chengjie Wang , Xiaobin Hu , Shuicheng Yan

Vision-Language Models (VLMs) have achieved strong results in video understanding, yet a key question remains: do they truly comprehend visual content or only learn shallow correlations between vision and language? Real visual…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Zongxia Li , Xiyang Wu , Guangyao Shi , Yubin Qin , Hongyang Du , Fuxiao Liu , Tianyi Zhou , Dinesh Manocha , Jordan Lee Boyd-Graber

Most of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The increased number of frames in long videos poses two main…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Ziyu Ma , Chenhui Gou , Hengcan Shi , Bin Sun , Shutao Li , Hamid Rezatofighi , Jianfei Cai

The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community transitions towards Reinforcement Learning (RL) and agentic…

Humans naturally share information with those they are connected to, and video has become one of the dominant mediums for communication and expression on the Internet. To support the creation of high-quality large-scale video content, a…

Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ailing Zhang , Lina Lei , Dehong Kong , Zhixin Wang , Jiaqi Xu , Fenglong Song , Chun-Le Guo , Chang Liu , Fan Li , Jie Chen

Anomaly analysis in surveillance videos is a crucial topic in computer vision. In recent years, multimodal large language models (MLLMs) have outperformed task-specific models in various domains. Although MLLMs are particularly versatile,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Haoran Chen , Dong Yi , Moyan Cao , Chensen Huang , Guibo Zhu , Jinqiao Wang

With the rapid advancement of Multimodal Large Language Models (MLLMs), numerous evaluation benchmarks have emerged. However, comprehensive assessments of their performance across diverse industrial applications remain limited. In this…

Computation and Language · Computer Science 2025-01-29 Dongyi Yi , Guibo Zhu , Chenglin Ding , Zongshu Li , Dong Yi , Jinqiao Wang

Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmentation (RVOS), it places emphasis on motion as well as…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Hao Fang , Runmin Cong , Xiankai Lu , Zhiyang Chen , Wei Zhang

Comprehending long videos remains a significant challenge for Large Multi-modal Models (LMMs). Current LMMs struggle to process even minutes to hours videos due to their lack of explicit memory and retrieval mechanisms. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Sameer Malik , Moyuru Yamada , Ayush Singh , Dishank Aggarwal

Existing video understanding benchmarks often conflate knowledge-based and purely image-based questions, rather than clearly isolating a model's temporal reasoning ability, which is the key aspect that distinguishes video understanding from…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Bo Feng , Zhengfeng Lai , Shiyu Li , Zizhen Wang , Simon Wang , Ping Huang , Meng Cao

Notable breakthroughs in unified understanding and generation modeling have led to remarkable advancements in image understanding, reasoning, production and editing, yet current foundational models predominantly focus on processing images,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Zhiyu Tan , Hao Yang , Luozheng Qin , Jia Gong , Mengping Yang , Hao Li

Large vision-language models (LVLMs) have recently achieved rapid progress, sparking numerous studies to evaluate their multi-modal capabilities. However, we dig into current evaluation works and identify two primary issues: 1) Visual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Lin Chen , Jinsong Li , Xiaoyi Dong , Pan Zhang , Yuhang Zang , Zehui Chen , Haodong Duan , Jiaqi Wang , Yu Qiao , Dahua Lin , Feng Zhao

Video understanding aims to enable models to perceive, reason about, and interact with the dynamic visual world. In contrast to image understanding, video understanding inherently requires modeling temporal dynamics and evolving visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Zhaochong An , Zirui Li , Mingqiao Ye , Feng Qiao , Jiaang Li , Zongwei Wu , Vishal Thengane , Chengzu Li , Lei Li , Luc Van Gool , Guolei Sun , Serge Belongie

Understanding visually-rich business documents to extract structured data and automate business workflows has been receiving attention both in academia and industry. Although recent multi-modal language models have achieved impressive…

Computation and Language · Computer Science 2023-09-19 Zilong Wang , Yichao Zhou , Wei Wei , Chen-Yu Lee , Sandeep Tata

Current video understanding models excel at recognizing "what" is happening but fall short in high-level cognitive tasks like causal reasoning and future prediction, a limitation rooted in their lack of commonsense world knowledge. To…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 L'ea Dubois , Klaus Schmidt , Chengyu Wang , Ji-Hoon Park , Lin Wang , Santiago Munoz

Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video durations. Mastering this intense cognitive load constitutes…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Haichen He , Jiayi Zhou , Sifeng Shang , Yihan Hu , Yuanhan Zhang , Kaiyang Zhou

Evaluating generative foundation models on open-ended multimodal understanding (MMU) and generation (MMG) tasks across diverse modalities (e.g., images, audio, video) poses significant challenges due to the complexity of cross-modal…

Computation and Language · Computer Science 2025-03-25 Shu Pu , Yaochen Wang , Dongping Chen , Yuhang Chen , Guohao Wang , Qi Qin , Zhongyi Zhang , Zhiyuan Zhang , Zetong Zhou , Shuang Gong , Yi Gui , Yao Wan , Philip S. Yu

With the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character recognition (Video OCR), has become a crucial capability. This…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Yulin Fei , Yuhui Gao , Xingyuan Xian , Xiaojin Zhang , Tao Wu , Wei Chen

Vision-language models (VLMs) exhibit strong zero-shot generalization on natural images and show early promise in interpretable medical image analysis. However, existing benchmarks do not systematically evaluate whether these models truly…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Tianhong Zhou , Yin Xu , Yingtao Zhu , Chuxi Xiao , Haiyang Bian , Lei Wei , Xuegong Zhang