中文
相关论文

相关论文: An Annotated Video Dataset for Computing Video Mem…

200 篇论文

Online video websites receive huge amount of videos daily from users all around the world. How to provide valuable recommendations to viewers is an important task for both video websites and related third parties, such as search engines.…

社会与信息网络 · 计算机科学 2013-12-30 Qingbo Hu , Guan Wang , Philip S. Yu

This paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage both consensus and disagreement among…

计算机视觉与模式识别 · 计算机科学 2019-09-05 Hang Zhao , Antonio Torralba , Lorenzo Torresani , Zhicheng Yan

Web video is often used as a source of data in various fields of study. While specialized subsets of web video, mainly earmarked for dedicated purposes, are often analyzed in detail, there is little information available about the…

多媒体 · 计算机科学 2017-07-06 Luca Rossetto , Heiko Schuldt

The 3rd annual installment of the ActivityNet Large- Scale Activity Recognition Challenge, held as a full-day workshop in CVPR 2018, focused on the recognition of daily life, high-level, goal-oriented activities from user-generated videos…

To understand the world, we humans constantly need to relate the present to the past, and put events in context. In this paper, we enable existing video models to do the same. We propose a long-term feature bank---supportive information…

计算机视觉与模式识别 · 计算机科学 2019-04-19 Chao-Yuan Wu , Christoph Feichtenhofer , Haoqi Fan , Kaiming He , Philipp Krähenbühl , Ross Girshick

The share of videos in the internet traffic has been growing, therefore understanding how videos capture attention on a global scale is also of growing importance. Most current research focus on modeling the number of views, but we argue…

社会与信息网络 · 计算机科学 2018-04-12 Siqi Wu , Marian-Andrei Rizoiu , Lexing Xie

Action anticipation and forecasting in videos do not require a hat-trick, as far as there are signs in the context to foresee how actions are going to be deployed. Capturing these signs is hard because the context includes the past. We…

计算机视觉与模式识别 · 计算机科学 2019-01-15 Fiora Pirri , Lorenzo Mauro , Edoardo Alati , Valsamis Ntouskos , Mahdieh Izadpanahkakhk , Elham Omrani

We explore how reconciling several foundation models (large language models and vision-language models) with a novel unified memory mechanism could tackle the challenging video understanding problem, especially capturing the long-term…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Yue Fan , Xiaojian Ma , Rujie Wu , Yuntao Du , Jiaqi Li , Zhi Gao , Qing Li

Video captioning which automatically translates video clips into natural language sentences is a very important task in computer vision. By virtue of recent deep learning technologies, e.g., convolutional neural networks (CNNs) and…

计算机视觉与模式识别 · 计算机科学 2016-11-18 Junbo Wang , Wei Wang , Yan Huang , Liang Wang , Tieniu Tan

Multimedia content, such as advertisements and story videos, exhibit a rich blend of creativity and multiple modalities. They incorporate elements like text, visuals, audio, and storytelling techniques, employing devices like emotions,…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Aanisha Bhattacharya , Yaman K Singla , Balaji Krishnamurthy , Rajiv Ratn Shah , Changyou Chen

The annotation of video tampering dataset is a boring task that takes a lot of manpower and financial resources. At present, there is no published literature which is capable to improve the annotation efficiency of forged videos. We…

多媒体 · 计算机科学 2018-02-08 Ye Yao

This paper introduces a novel activity dataset which exhibits real-life and diverse scenarios of complex, temporally-extended human activities and actions. The dataset presents a set of videos of actors performing everyday activities in a…

计算机视觉与模式识别 · 计算机科学 2017-09-22 Jawad Tayyub , Majd Hawasly , David C. Hogg , Anthony G. Cohn

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Chaoyu Li , Sid Padmanabhuni , Maryam Cheema , Hasti Seifi , Pooyan Fazli

In this paper, we introduce the FOCAL (Ford-OLIVES Collaboration on Active Learning) dataset which enables the study of the impact of annotation-cost within a video active learning setting. Annotation-cost refers to the time it takes an…

计算机视觉与模式识别 · 计算机科学 2023-11-20 Kiran Kokilepersaud , Yash-Yee Logan , Ryan Benkert , Chen Zhou , Mohit Prabhushankar , Ghassan AlRegib , Enrique Corona , Kunjan Singh , Mostafa Parchami

With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limitation of current training datasets is their reliance on…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Xiaoyu Lin , Aniket Ghorpade , Hansheng Zhu , Justin Qiu , Dea Rrozhani , Monica Lama , Mick Yang , Zixuan Bian , Ruohan Ren , Alan B. Hong , Jiatao Gu , Chris Callison-Burch

This dissertation presents a methodology for recording speed climbing training sessions with multiple cameras and annotating the videos with relevant data, including body position, hand and foot placement, and timing. The annotated data is…

计算机视觉与模式识别 · 计算机科学 2023-05-24 Yufei Xie , Shaoman Li , Penghui Lin

Inferring physical actions from visual observations is a fundamental capability for advancing machine intelligence in the physical world. Achieving this requires large-scale, open-vocabulary video action datasets that span broad domains. We…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Delong Chen , Tejaswi Kasarla , Yejin Bang , Mustafa Shukor , Willy Chung , Jade Yu , Allen Bolourchi , Theo Moutakanni , Pascale Fung

While many action recognition datasets consist of collections of brief, trimmed videos each containing a relevant action, videos in the real-world (e.g., on YouTube) exhibit very different properties: they are often several minutes long,…

计算机视觉与模式识别 · 计算机科学 2019-09-02 Bruno Korbar , Du Tran , Lorenzo Torresani

This paper introduces a new challenge and datasets to foster research toward designing systems that can understand medical videos and provide visual answers to natural language questions. We believe medical videos may provide the best…

计算机视觉与模式识别 · 计算机科学 2022-02-01 Deepak Gupta , Kush Attal , Dina Demner-Fushman

Can out-of-the-box pretrained Large Language Models (LLMs) detect human affect successfully when observing a video? To address this question, for the first time, we evaluate comprehensively the capacity of popular LLMs for successfully…

计算机视觉与模式识别 · 计算机科学 2026-01-30 David Melhart , Matthew Barthet , Georgios N. Yannakakis