中文
相关论文

相关论文: RefineShot: Rethinking Cinematography Understandin…

200 篇论文

We introduce MoviePuzzle, a novel challenge that targets visual narrative reasoning and holistic movie understanding. Despite the notable progress that has been witnessed in the realm of video understanding, most prior works fail to present…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Jianghui Wang , Yuxuan Wang , Dongyan Zhao , Zilong Zheng

Instruction-driven image editing with unified multimodal generative models has advanced rapidly, yet their underlying visual reasoning remains limited, leading to suboptimal performance on reasoning-centric edits. Reinforcement learning…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Hengjia Li , Liming Jiang , Qing Yan , Yizhi Song , Hao Kang , Zichuan Liu , Xin Lu , Boxi Wu , Deng Cai

Movie screenplays are rich long-form narratives that interleave complex character relationships, temporally ordered events, and dialogue-driven interactions. While prior benchmarks target individual subtasks such as question answering or…

计算与语言 · 计算机科学 2026-05-27 Qiuyu Tian , Zequn Liu , Yiding Li , Fengyi Chen , Youyong Kong , Fan Guo , Yuyao Li , Jinjing Shen , Zhijing Xie , Yiyun Luo , Xin Zhang , Yingce Xia

Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ailing Zhang , Lina Lei , Dehong Kong , Zhixin Wang , Jiaqi Xu , Fenglong Song , Chun-Le Guo , Chang Liu , Fan Li , Jie Chen

The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consistency. However, despite recent breakthroughs such as Veo…

We propose RocketScience, an open-source contrastive VLM benchmark that tests for spatial relation understanding. It is comprised of entirely new real-world image-text pairs covering mostly relative spatial understanding and the order of…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Nils Hoehing , Mayug Maniparambil , Ellen Rushe , Noel E. O'Connor , Anthony Ventresque

Story visualization aims to generate coherent image sequences that faithfully represent a narrative and match given character references. Despite progress in generative models, existing benchmarks remain narrow in scope, often limited to…

Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existing evaluation protocols are limited to ambiguous language modeling tasks. To address this…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Vladimir Zaigrajew , Dawid Pludowski , Hubert Baniecki , Przemyslaw Biecek

Geometric problem solving constitutes a critical branch of mathematical reasoning, requiring precise analysis of shapes and spatial relationships. Current evaluations of geometric reasoning in vision-language models (VLMs) face limitations,…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Yuan Feng , Yue Yang , Xiaohan He , Jiatong Zhao , Jianlong Chen , Zijun Chen , Daocheng Fu , Qi Liu , Renqiu Xia , Bo Zhang , Junchi Yan

Visual understanding goes well beyond object recognition. With one glance at an image, we can effortlessly imagine the world beyond the pixels: for instance, we can infer people's actions, goals, and mental states. While this task is easy…

计算机视觉与模式识别 · 计算机科学 2019-03-27 Rowan Zellers , Yonatan Bisk , Ali Farhadi , Yejin Choi

Image captioning has long been a pivotal task in visual understanding, with recent advancements in vision-language models (VLMs) significantly enhancing the ability to generate detailed image captions. However, the evaluation of detailed…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Qinghao Ye , Xianhan Zeng , Fu Li , Chunyuan Li , Haoqi Fan

Reliable evaluation benchmarks designed for replicability and comprehensiveness have driven progress in machine learning. Due to the lack of a multilingual benchmark, however, vision-and-language research has mostly focused on English…

计算与语言 · 计算机科学 2022-07-19 Emanuele Bugliarello , Fangyu Liu , Jonas Pfeiffer , Siva Reddy , Desmond Elliott , Edoardo Maria Ponti , Ivan Vulić

Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Silin Gao , Sheryl Mathew , Li Mi , Sepideh Mamooler , Mengjie Zhao , Hiromi Wakaki , Yuki Mitsufuji , Syrielle Montariol , Antoine Bosselut

Automated surgical workflow analysis is crucial for education, research, and clinical decision-making, but the lack of annotated datasets hinders the development of accurate and comprehensive workflow analysis solutions. We introduce a…

计算机视觉与模式识别 · 计算机科学 2025-03-17 David Gastager , Ghazal Ghazaei , Constantin Patsch

Multimodal large language models (MLLMs) have demonstrated powerful capabilities in general spatial understanding and reasoning. However, their fine-grained spatial understanding and reasoning capabilities in complex urban scenarios have…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Jun Zhang , Jie Feng , Long Chen , Junhui Wang , Zhicheng Liu , Depeng Jin , Yong Li

Video is an increasingly prominent and information-dense medium, yet it poses substantial challenges for language models. A typical video consists of a sequence of shorter segments, or shots, that collectively form a coherent narrative.…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Richard Luo , Austin Peng , Adithya Vasudev , Rishabh Jain

Visual Language Models (VLMs) are now sufficiently advanced to support a broad range of applications, including answering complex visual questions, and are increasingly expected to interact with images in varied ways. To evaluate them,…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Ludovic Arnould , Salim Khazem , Hugues Ali Mehenni

While recent advancements in vision-language models have improved video understanding, diagnosing their capacity for deep, narrative comprehension remains a challenge. Existing benchmarks often test short-clip recognition or use…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Nisarg A. Shah , Amir Ziai , Chaitanya Ekanadham , Vishal M. Patel

Real world deployments often expose modern object recognition models to domain shifts that precipitate a severe drop in accuracy. Such shifts encompass (i) variations in low level image statistics, (ii) changes in object pose and viewpoint,…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Junghyun Park , Tuan Anh Nguyen , Dugki Min

The few-shot natural language understanding (NLU) task has attracted much recent attention. However, prior methods have been evaluated under a disparate set of protocols, which hinders fair comparison and measuring progress of the field. To…