中文
相关论文

相关论文: Building Scalable Video Understanding Benchmarks t…

200 篇论文

We introduce \textbf{LongInsightBench}, the first benchmark designed to assess models' ability to understand long videos, with a focus on human language, viewpoints, actions, and other contextual elements, while integrating \textbf{visual,…

计算机视觉与模式识别 · 计算机科学 2025-10-22 ZhaoYang Han , Qihan Lin , Hao Liang , Bowen Chen , Zhou Liu , Wentao Zhang

The task of action spotting consists in both identifying actions and precisely localizing them in time with a single timestamp in long, untrimmed video streams. Automatically extracting those actions is crucial for many sports applications,…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Silvio Giancola , Anthony Cioppa , Bernard Ghanem , Marc Van Droogenbroeck

Understanding sports is crucial for the advancement of Natural Language Processing (NLP) due to its intricate and dynamic nature. Reasoning over complex sports scenarios has posed significant challenges to current NLP technologies which…

计算与语言 · 计算机科学 2024-06-24 Zhengbang Yang , Haotian Xia , Jingxi Li , Zezhi Chen , Zhuangdi Zhu , Weining Shen

Automatic summarization generation of sports video content has been object of great interest for many years. Although semantic descriptions techniques have been proposed, many of the approaches still rely on low-level video descriptors that…

信息检索 · 计算机科学 2014-11-25 Arnau Raventos , Raul Quijada , Luis Torres , Francesc Tarres

This work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Jiahao Shao , Yuanbo Yang , Hongyu Zhou , Youmin Zhang , Yujun Shen , Vitor Guizilini , Yue Wang , Matteo Poggi , Yiyi Liao

In this paper, we introduce VCSL (Video Copy Segment Localization), a new comprehensive segment-level annotated video copy dataset. Compared with existing copy detection datasets restricted by either video-level annotation or small-scale,…

计算机视觉与模式识别 · 计算机科学 2022-06-17 Sifeng He , Xudong Yang , Chen Jiang , Gang Liang , Wei Zhang , Tan Pan , Qing Wang , Furong Xu , Chunguang Li , Jingxiong Liu , Hui Xu , Kaiming Huang , Yuan Cheng , Feng Qian , Xiaobo Zhang , Lei Yang

Tennis is one of the most widely followed sports, generating extensive broadcast footage with strong potential for professional analysis, automated coaching, and real-time commentary. However, automatic tennis understanding remains…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Zhaoyu Liu , Xi Weng , Lianyu Hu , Zhe Hou , Kan Jiang , Jin Song Dong , Yang Liu

Recent progress in multi-modal large language models (MLLMs) has significantly advanced video understanding. However, their performance on long-form videos remains limited by computational constraints and suboptimal frame selection. We…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Wenhui Tan , Ruihua Song , Jiaze Li , Jianzhong Ju , Zhenbo Luo

Video anomaly detection research is generally evaluated on short, isolated benchmark videos only a few minutes long. However, in real-world environments, security cameras observe the same scene for months or years at a time, and the notion…

计算机视觉与模式识别 · 计算机科学 2024-04-12 Zhengye Yang , Richard Radke

We propose a novel supervised learning technique for summarizing videos by automatically selecting keyframes or key subshots. Casting the problem as a structured prediction problem on sequential data, our main idea is to use Long Short-Term…

计算机视觉与模式识别 · 计算机科学 2016-08-01 Ke Zhang , Wei-Lun Chao , Fei Sha , Kristen Grauman

Sports video data is recorded for nearly every major tournament but remains archived and inaccessible to large scale data mining and analytics. It can only be viewed sequentially or manually tagged with higher-level labels which is time…

计算机视觉与模式识别 · 计算机科学 2017-12-27 Anurag Ghosh , Suriya Singh , C. V. Jawahar

Contrastive Language-Image Pre-training (CLIP) has been the cornerstone for zero-shot classification, text-image retrieval, and text-image generation by aligning image and text modalities. Despite its widespread adoption, a significant…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Beichen Zhang , Pan Zhang , Xiaoyi Dong , Yuhang Zang , Jiaqi Wang

With the rapid growth of video data on the internet, video summarization is becoming a very important AI technology. However, due to the high labelling cost of video summarization, existing studies have to be conducted on small-scale…

多媒体 · 计算机科学 2026-01-13 Cairong Zhao , Chutian Wang , Zifan Song , Guosheng Hu , Haonan Chen , Xiaofan Zhai

The objective of this paper is a temporal alignment network that ingests long term video sequences, and associated text sentences, in order to: (1) determine if a sentence is alignable with the video; and (2) if it is alignable, then…

计算机视觉与模式识别 · 计算机科学 2022-04-07 Tengda Han , Weidi Xie , Andrew Zisserman

Despite recent progress on the short-video Text-Visual Question Answering (ViteVQA) task - largely driven by benchmarks such as M4-ViteVQA - existing datasets still suffer from limited video duration and narrow evaluation scopes, making it…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Yangyang Zhong , Ji Qi , Yuan Yao , Pengxin Luo , Yunfeng Yan , Donglian Qi , Zhiyuan Liu , Tat-Seng Chua

Large-scale annotated datasets allow AI systems to learn from and build upon the knowledge of the crowd. Many crowdsourcing techniques have been developed for collecting image annotations. These techniques often implicitly rely on the fact…

人机交互 · 计算机科学 2016-10-07 Gunnar A. Sigurdsson , Olga Russakovsky , Ali Farhadi , Ivan Laptev , Abhinav Gupta

Video summarization is a technique to create a short skim of the original video while preserving the main stories/content. There exists a substantial interest in automatizing this process due to the rapid growth of the available material.…

计算机视觉与模式识别 · 计算机科学 2019-04-12 Mayu Otani , Yuta Nakashima , Esa Rahtu , Janne Heikkilä

Camera calibration and localization, sometimes simply named camera calibration, enables many applications in the context of soccer broadcasting, for instance regarding the interpretation and analysis of the game, or the insertion of…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Floriane Magera , Thomas Hoyoux , Olivier Barnich , Marc Van Droogenbroeck

Modeling fine-grained speaking styles remains challenging for language-speech representation pre-training, as existing speech-text models are typically trained with coarse captions or task-specific supervision, and scalable fine-grained…

音频与语音处理 · 电气工程与系统科学 2026-04-21 Yifan Yang , Bing Han , Hui Wang , Wei Wang , Ziyang Ma , Long Zhou , Zengrui Jin , Guanrou Yang , Tianrui Wang , Xu Tan , Xie Chen

Existing long video retrieval systems are trained and tested in the paragraph-to-video retrieval regime, where every long video is described by a single long paragraph. This neglects the richness and variety of possible valid descriptions…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Matthew Gwilliam , Michael Cogswell , Meng Ye , Karan Sikka , Abhinav Shrivastava , Ajay Divakaran