中文
相关论文

相关论文: MME-CoF-Pro: Evaluating Reasoning Coherence in Vid…

200 篇论文

Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism has enhanced reasoning in text-based tasks, its application…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Mi Luo , Zihui Xue , Alex Dimakis , Kristen Grauman

Recent advances in CoT reasoning and RL post-training have been reported to enhance video reasoning capabilities of MLLMs. This progress naturally raises a question: can these models perform complex video reasoning in a manner comparable to…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Junhao Cheng , Yuying Ge , Teng Wang , Yixiao Ge , Jing Liao , Ying Shan

Recent advances in video generation have posed great challenges in the assessment of AI-generated content, particularly with the emergence of increasingly sophisticated models. The various inconsistencies and defects observed in such videos…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Rui Chen , Lei Sun , Jing Tang , Geng Li , Xiangxiang Chu

Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true multimodal reasoning? Existing benchmarks fail to address this question rigorously, as…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Xiaotian Zhang , Jianhui Wei , Yuan Wang , Jie Tan , Yichen Li , Yan Zhang , Ziyi Chen , Daoan Zhang , Dezhi YU , Wei Xu , Songtao Jiang , Zuozhu Liu

The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consistency. However, despite recent breakthroughs such as Veo…

Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabilities. Prior work attributes this to a Chain-of-Frames (CoF) mechanism, where reasoning is assumed…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Ruisi Wang , Zhongang Cai , Fanyi Pu , Junxiang Xu , Wanqi Yin , Maijunxian Wang , Ran Ji , Chenyang Gu , Bo Li , Ziqi Huang , Hokin Deng , Dahua Lin , Ziwei Liu , Lei Yang

The advancement of Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs) and large vision-language models (LVLMs). However, a rigorous evaluation framework for video CoT reasoning…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Yukun Qi , Yiming Zhao , Yu Zeng , Xikun Bao , Wenxuan Huang , Lin Chen , Zehui Chen , Jie Zhao , Zhongang Qi , Feng Zhao

Recent work has shown that eliciting Large Language Models (LLMs) to generate reasoning traces in natural language before answering the user's request can significantly improve their performance across tasks. This approach has been extended…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Sara Ghazanfari , Francesco Croce , Nicolas Flammarion , Prashanth Krishnamurthy , Farshad Khorrami , Siddharth Garg

The ability to reason about temporal and causal events from videos lies at the core of human intelligence. Most video reasoning benchmarks, however, focus on pattern recognition from complex visual and language input, instead of on causal…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Kexin Yi , Chuang Gan , Yunzhu Li , Pushmeet Kohli , Jiajun Wu , Antonio Torralba , Joshua B. Tenenbaum

Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic…

Composed Video Retrieval (CoVR) aims to find a target video given a reference video and a textual modification. Prior work assumes the modification text fully specifies the visual changes, overlooking after-effects and implicit consequences…

Chain-of-thought (CoT) reasoning has been highly successful in solving complex tasks in natural language processing, and recent multimodal large language models (MLLMs) have extended this paradigm to video reasoning. However, these models…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Yiwu Zhong , Zi-Yuan Hu , Yin Li , Liwei Wang

Evaluating object removal in images and videos remains challenging because the task is inherently one-to-many, yet existing metrics frequently disagree with human perception. Full-reference metrics reward copy-paste behaviors over genuine…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Fuhao Li , Shaofeng You , Jiagao Hu , Yu Liu , Yuxuan Chen , Zepeng Wang , Fei Wang , Daiguo Zhou , Jian Luan

Despite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions. First, they overwhelmingly rely on static images, failing to capture the temporal complexity of…

Recent studies have demonstrated that incorporating Chain-of-Thought (CoT) reasoning into the detection process can enhance a model's ability to detect synthetic images. However, excessively lengthy reasoning incurs substantial resource…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Changjiang Jiang , Xinkuan Sha , Fengchang Yu , Jingjing Liu , Jian Liu , Mingqi Fang , Chenfeng Zhang , Wei Lu

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able…

Vision-language models (VLMs) have recently demonstrated strong efficacy as visual assistants that can parse natural queries about the visual content and generate human-like outputs. In this work, we explore the ability of these models to…

计算与语言 · 计算机科学 2024-03-21 Yangyi Chen , Karan Sikka , Michael Cogswell , Heng Ji , Ajay Divakaran

Reasoning over dynamic visual content remains a central challenge for multimodal large language models. Recent thinking models generate explicit reasoning traces for interpretability; however, their reasoning often appears convincing while…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Muhammad Maaz , Hanoona Rasheed , Fahad Shahbaz Khan , Salman Khan

Chain-of-thought (CoT) reasoning has emerged as a powerful tool for multimodal large language models on video understanding tasks. However, its necessity and advantages over direct answering remain underexplored. In this paper, we first…

Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories:…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Hui Han , Siyuan Li , Jiaqi Chen , Yiwen Yuan , Yuling Wu , Chak Tou Leong , Hanwen Du , Junchen Fu , Youhua Li , Jie Zhang , Chi Zhang , Li-jia Li , Yongxin Ni