中文
相关论文

相关论文: An Annotated Video Dataset for Computing Video Mem…

200 篇论文

We live in a world filled with never-ending streams of multimodal information. As a more natural recording of the real scenario, long form audio-visual videos are expected as an important bridge for better exploring and understanding the…

多媒体 · 计算机科学 2023-06-19 Wenxuan Hou , Guangyao Li , Yapeng Tian , Di Hu

We introduce a unified framework for generic video annotation with bounding boxes. Video annotation is a longstanding problem, as it is a tedious and time-consuming process. We tackle two important challenges of video annotation: (1)…

计算机视觉与模式识别 · 计算机科学 2020-12-24 A. Kuznetsova , A. Talati , Y. Luo , K. Simmons , V. Ferrari

The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more…

We present an approach to labeling short video clips with English verbs as event descriptions. A key distinguishing aspect of this work is that it labels videos with verbs that describe the spatiotemporal interaction between event…

Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Weihan Wang , Zehai He , Wenyi Hong , Yean Cheng , Xiaohan Zhang , Ji Qi , Xiaotao Gu , Shiyu Huang , Bin Xu , Yuxiao Dong , Ming Ding , Jie Tang

Video summarization aims to distill the most important information from a source video to produce either an abridged clip or a textual narrative. Traditionally, different methods have been proposed depending on whether the output is a video…

计算机视觉与模式识别 · 计算机科学 2024-04-24 Jingyang Lin , Hang Hua , Ming Chen , Yikang Li , Jenhao Hsiao , Chiuman Ho , Jiebo Luo

Manual spatio-temporal annotation of human action in videos is laborious, requires several annotators and contains human biases. In this paper, we present a weakly supervised approach to automatically obtain spatio-temporal annotations of…

计算机视觉与模式识别 · 计算机科学 2016-05-27 Waqas Sultani , Mubarak Shah

We present a new model to determine relative skill from long videos, through learnable temporal attention modules. Skill determination is formulated as a ranking problem, making it suitable for common and generic tasks. However, for long…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Hazel Doughty , Walterio Mayol-Cuevas , Dima Damen

Retrieving target videos based on text descriptions is a task of great practical value and has received increasing attention over the past few years. Despite recent progress, imperfect annotations in existing video retrieval datasets have…

计算机视觉与模式识别 · 计算机科学 2022-07-22 Zeyu Wang , Yu Wu , Karthik Narasimhan , Olga Russakovsky

Memorability determines what evanesces into emptiness, and what worms its way into the deepest furrows of our minds. It is the key to curating more meaningful media content as we wade through daily digital torrents. The Predicting Media…

多媒体 · 计算机科学 2021-01-01 Lorin Sweeney , Graham Healy , Alan F. Smeaton

Manual annotation remains the gold standard for high-quality, dense temporal video datasets, yet it is inherently time-consuming. Vision-language models can aid human annotators and expedite this process. We report on the impact of…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Juan Gutiérrez , Victor Gutiérrez , Ángel Mora , Silvia Rodriguez , José Luis Blanco

This paper describes the AVA-Kinetics localized human actions video dataset. The dataset is collected by annotating videos from the Kinetics-700 dataset using the AVA annotation protocol, and extending the original AVA dataset with these…

计算机视觉与模式识别 · 计算机科学 2020-05-21 Ang Li , Meghana Thotakuri , David A. Ross , João Carreira , Alexander Vostrikov , Andrew Zisserman

This paper introduces a new large consent-driven dataset aimed at assisting in the evaluation of algorithmic bias and robustness of computer vision and audio speech models in regards to 11 attributes that are self-provided or labeled by…

计算机视觉与模式识别 · 计算机科学 2023-03-10 Bilal Porgali , Vítor Albiero , Jordan Ryda , Cristian Canton Ferrer , Caner Hazirbas

Micro-videos are six-second videos popular on social media networks with several unique properties. Firstly, because of the authoring process, they contain significantly more diversity and narrative structure than existing collections of…

计算机视觉与模式识别 · 计算机科学 2016-04-04 Phuc Xuan Nguyen , Gregory Rogez , Charless Fowlkes , Deva Ramanan

Humans have a selective memory, remembering relevant episodes and forgetting the less relevant information. Possessing awareness of event memorability for a user could help intelligent systems in more accurate user modelling, especially for…

人机交互 · 计算机科学 2025-07-21 Maria Tsfasman , Ramin Ghorbani , Catholijn M. Jonker , Bernd Dudzik

High-quality and consistent annotations are fundamental to the successful development of robust machine learning models. Traditional data annotation methods are resource-intensive and inefficient, often leading to a reliance on third-party…

计算机视觉与模式识别 · 计算机科学 2024-02-12 Amir Ziai , Aneesh Vartakavi

Video summarization is a technique to create a short skim of the original video while preserving the main stories/content. There exists a substantial interest in automatizing this process due to the rapid growth of the available material.…

计算机视觉与模式识别 · 计算机科学 2019-04-12 Mayu Otani , Yuta Nakashima , Esa Rahtu , Janne Heikkilä

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activity analysis, and…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Wenqi Zhou , Kai Cao , Hao Zheng , Yunze Liu , Xinyi Zheng , Miao Liu , Per Ola Kristensson , Walterio Mayol-Cuevas , Fan Zhang , Weizhe Lin , Junxiao Shen

We introduce Replay, a collection of multi-view, multi-modal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with several static cameras, as well as wearable action cameras,…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Roman Shapovalov , Yanir Kleiman , Ignacio Rocco , David Novotny , Andrea Vedaldi , Changan Chen , Filippos Kokkinos , Ben Graham , Natalia Neverova

Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Jingyang Lin , Jialian Wu , Ximeng Sun , Ze Wang , Jiang Liu , Yusheng Su , Xiaodong Yu , Hao Chen , Jiebo Luo , Zicheng Liu , Emad Barsoum