English
Related papers

Related papers: VCBench: A Streaming Counting Benchmark for Spatia…

200 papers

An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos, predict rich, detailed textual descriptions, and be able to produce outputs before processing…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Xingyi Zhou , Anurag Arnab , Shyamal Buch , Shen Yan , Austin Myers , Xuehan Xiong , Arsha Nagrani , Cordelia Schmid

Video generation has advanced significantly, evolving from producing unrealistic outputs to generating videos that appear visually convincing and temporally coherent. To evaluate these video generative models, benchmarks such as VBench have…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Dian Zheng , Ziqi Huang , Hongbo Liu , Kai Zou , Yinan He , Fan Zhang , Lulu Gu , Yuanhan Zhang , Jingwen He , Wei-Shi Zheng , Yu Qiao , Ziwei Liu

Video Large Multimodal Models (VLMMs) have shown impressive performance in video understanding, yet their ability to accurately capture the temporal order of multiple events remains underexplored. We interestingly observe that, even when…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Daechul Ahn , Yura Choi , Hyeonbeom Choi , Seongwon Cho , San Kim , Jonghyun Choi

Evaluating the nuanced human-centric video understanding capabilities of Multimodal Large Language Models (MLLMs) remains a great challenge, as existing benchmarks often overlook the intricacies of emotion, behavior, and cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Ting Zhou , Daoyuan Chen , Qirui Jiao , Bolin Ding , Yaliang Li , Ying Shen

Recent advances in creative AI have enabled the synthesis of high-fidelity images and videos conditioned on language instructions. Building on these developments, text-to-video diffusion models have evolved into embodied world models (EWMs)…

Robotics · Computer Science 2025-05-20 Hu Yue , Siyuan Huang , Yue Liao , Shengcong Chen , Pengfei Zhou , Liliang Chen , Maoqing Yao , Guanghui Ren

Video object segmentation (VOS) aims to distinguish and track target objects in a video. Despite the excellent performance achieved by off-the-shell VOS models, existing VOS benchmarks mainly focus on short-term videos lasting about 5…

Computer Vision and Pattern Recognition · Computer Science 2024-05-02 Lingyi Hong , Zhongying Liu , Wenchao Chen , Chenzhi Tan , Yuang Feng , Xinyu Zhou , Pinxue Guo , Jinglun Li , Zhaoyu Chen , Shuyong Gao , Wei Zhang , Wenqiang Zhang

How (dis)similar are the learning trajectories of vision-language models and children? Recent modeling work has attempted to understand the gap between models' and humans' data efficiency by constructing models trained on less data,…

Computation and Language · Computer Science 2024-12-10 Alvin Wei Ming Tan , Sunny Yu , Bria Long , Wanjing Anya Ma , Tonya Murray , Rebecca D. Silverman , Jason D. Yeatman , Michael C. Frank

We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been proposed to pursue this…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Xiang Li , Jian Ding , Mohamed Elhoseiny

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic consistency and factual accuracy remains largely understudied.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Yubin Chen , Xuyang Guo , Zhenmei Shi , Zhao Song , Jiahao Zhang

Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the lack of relevant benchmarks. Previous work in visual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yiqing Shen , Chenjia Li , Chenxiao Fan , Mathias Unberath

With the rapid advancement of video understanding, existing benchmarks are becoming increasingly saturated, exposing a critical discrepancy between inflated leaderboard scores and real-world model capabilities. To address this widening gap,…

Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insufficient. To address this, we introduce InstructionBench, an…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Haiwan Wei , Yitian Yuan , Xiaohan Lan , Wei Ke , Lin Ma

Recent advancements in video generation models, like Stable Video Diffusion, show promising results, but primarily focus on short, single-scene videos. These models struggle with generating long videos that involve multiple scenes, coherent…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Weijia Wu , Mingyu Liu , Zeyu Zhu , Xi Xia , Haoen Feng , Wen Wang , Kevin Qinghong Lin , Chunhua Shen , Mike Zheng Shou

Recent advances in video generation models demonstrate their potential as world simulators, but they often struggle with videos deviating from physical laws, a key concern overlooked by most text-to-video benchmarks. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Yongfan Chen , Xiuwen Zhu , Tianyu Li

Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hours) -- existing benchmarks either neglect this multi-scale…

Open-ended video game glitch detection aims to identify glitches in gameplay videos, describe them in natural language, and localize when they occur. Unlike conventional game glitch understanding tasks which have largely been framed as…

Multiagent Systems · Computer Science 2026-04-27 Muyang Zheng , Tong Zhou , Geyang Wu , Zihao Lin , Haibo Wang , Lifu Huang

Video event extraction aims to detect salient events from a video and identify the arguments for each event as well as their semantic roles. Existing methods focus on capturing the overall visual scene of each frame, ignoring fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Guang Yang , Manling Li , Jiajie Zhang , Xudong Lin , Shih-Fu Chang , Heng Ji

Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos due to large viewpoint variations, rapid shot transitions, and cluttered scenes, it…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Ismael Elsharkawi , Ahmed Sait , Silvio Giancola , Bernard Ghanem , Hossam Sharara , Abdelrahman Eldesokey

Multiple existing benchmarks involve tracking and segmenting objects in video e.g., Video Object Segmentation (VOS) and Multi-Object Tracking and Segmentation (MOTS), but there is little interaction between them due to the use of disparate…

Computer Vision and Pattern Recognition · Computer Science 2022-11-23 Ali Athar , Jonathon Luiten , Paul Voigtlaender , Tarasha Khurana , Achal Dave , Bastian Leibe , Deva Ramanan

Automatic identification of events and recurrent behavior analysis are critical for video surveillance. However, most existing content-based video retrieval benchmarks focus on scene-level similarity and do not evaluate the action…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Oriol Rabasseda , Zenjie Li , Kamal Nasrollahi , Sergio Escalera
‹ Prev 1 4 5 6 7 8 10 Next ›