中文
相关论文

相关论文: Dense-Captioning Events in Videos

200 篇论文

In this paper, we introduce SoccerNet, a benchmark for action spotting in soccer videos. The dataset is composed of 500 complete soccer games from six main European leagues, covering three seasons from 2014 to 2017 and a total duration of…

计算机视觉与模式识别 · 计算机科学 2019-03-26 Silvio Giancola , Mohieddine Amine , Tarek Dghaily , Bernard Ghanem

In this work\footnote {This work was supported in part by the National Science Foundation under grant IIS-1212948.}, we present a method to represent a video with a sequence of words, and learn the temporal sequencing of such words as the…

计算机视觉与模式识别 · 计算机科学 2019-06-18 Sangwoo Cho , Hassan Foroosh

Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes. A complex event relates to dense crowds, anomalous individuals, or…

计算机视觉与模式识别 · 计算机科学 2023-07-14 Weiyao Lin , Huabin Liu , Shizhan Liu , Yuxi Li , Rui Qian , Tao Wang , Ning Xu , Hongkai Xiong , Guo-Jun Qi , Nicu Sebe

Action parsing in videos with complex scenes is an interesting but challenging task in computer vision. In this paper, we propose a generic 3D convolutional neural network in a multi-task learning manner for effective Deep Action Parsing…

计算机视觉与模式识别 · 计算机科学 2016-02-11 Li Liu , Yi Zhou , Ling Shao

A short clip of video may contain progression of multiple events and an interesting story line. A human need to capture both the event in every shot and associate them together to understand the story behind it. In this work, we present a…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Mingfei Han , Linjie Yang , Xiaojun Chang , Lina Yao , Heng Wang

This notebook paper presents an overview and comparative analysis of our systems designed for the following three tasks in ActivityNet Challenge 2019: trimmed action recognition, dense-captioning events in videos, and spatio-temporal action…

计算机视觉与模式识别 · 计算机科学 2019-06-18 Zhaofan Qiu , Dong Li , Yehao Li , Qi Cai , Yingwei Pan , Ting Yao

This work aims to present novel description methods for human action recognition. Generally, a video sequence can be represented as a collection of spatial temporal words by detecting space-time interest points and describing the unique…

人机交互 · 计算机科学 2011-01-04 Ruoyun Gao , Michael S. Lew , Ling Shao

Counting repetitive actions are widely seen in human activities such as physical exercise. Existing methods focus on performing repetitive action counting in short videos, which is tough for dealing with longer videos in more realistic…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Huazhang Hu , Sixun Dong , Yiqun Zhao , Dongze Lian , Zhengxin Li , Shenghua Gao

Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the {sparsity dilemma} in video annotations, which fails to provide the context information between…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Hongxiang Li , Meng Cao , Xuxin Cheng , Zhihong Zhu , Yaowei Li , Yuexian Zou

Egocentric vision consists in acquiring images along the day from a first person point-of-view using wearable cameras. The automatic analysis of this information allows to discover daily patterns for improving the quality of life of the…

计算机视觉与模式识别 · 计算机科学 2017-11-10 Marc Bolaños , Álvaro Peris , Francisco Casacuberta , Sergi Soler , Petia Radeva

This paper presents a new method to describe spatio-temporal relations between objects and hands, to recognize both interactions and activities within video demonstrations of manual tasks. The approach exploits Scene Graphs to extract key…

计算机视觉与模式识别 · 计算机科学 2023-07-10 Elena Merlo , Marta Lagomarsino , Edoardo Lamon , Arash Ajoudani

This paper aims for event recognition when video examples are scarce or even completely absent. The key in such a challenging setting is a semantic video representation. Rather than building the representation from individual attribute…

计算机视觉与模式识别 · 计算机科学 2015-11-10 Amirhossein Habibian , Thomas Mensink , Cees G. M. Snoek

Video action recognition is one of the representative tasks for video understanding. Over the last decade, we have witnessed great advancements in video action recognition thanks to the emergence of deep learning. But we also encountered…

计算机视觉与模式识别 · 计算机科学 2020-12-14 Yi Zhu , Xinyu Li , Chunhui Liu , Mohammadreza Zolfaghari , Yuanjun Xiong , Chongruo Wu , Zhi Zhang , Joseph Tighe , R. Manmatha , Mu Li

Segmenting video content into events provides semantic structures for indexing, retrieval, and summarization. Since motion cues are not available in continuous photo-streams, and annotations in lifelogging are scarce and costly, the frames…

计算机视觉与模式识别 · 计算机科学 2018-08-08 Ana Garcia del Molino , Joo-Hwee Lim , Ah-Hwee Tan

Video captioning generate a sentence that describes the video content. Existing methods always require a number of captions (\eg, 10 or 20) per video to train the model, which is quite costly. In this work, we explore the possibility of…

计算机视觉与模式识别 · 计算机科学 2024-11-07 Ping Li , Tao Wang , Xinkui Zhao , Xianghua Xu , Mingli Song

Natural language descriptions of user interface (UI) elements such as alternative text are crucial for accessibility and language-based interaction in general. Yet, these descriptions are constantly missing in mobile UIs. We propose widget…

机器学习 · 计算机科学 2020-10-12 Yang Li , Gang Li , Luheng He , Jingjie Zheng , Hong Li , Zhiwei Guan

In video understanding, action spotting consists in temporally localizing human-induced events annotated with single timestamps. In this paper, we propose a novel loss function that specifically considers the temporal context naturally…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Anthony Cioppa , Adrien Deliège , Silvio Giancola , Bernard Ghanem , Marc Van Droogenbroeck , Rikke Gade , Thomas B. Moeslund

Dense video captioning aims to temporally localize events in video and generate captions for each event. While recent works propose end-to-end models, they suffer from two limitations: (1) applying timestamp supervision only to text while…

计算机视觉与模式识别 · 计算机科学 2025-09-08 MinJu Jeon , Si-Woo Kim , Ye-Chan Kim , HyunGee Kim , Dong-Jin Kim

The proliferation of video content on platforms like YouTube and Vimeo presents significant challenges in efficiently locating relevant information. Automatic video summarization aims to address this by extracting and presenting key content…

计算机视觉与模式识别 · 计算机科学 2024-08-28 Jia-Hong Huang

This paper proposes a network architecture to perform variable length semantic video generation using captions. We adopt a new perspective towards video generation where we allow the captions to be combined with the long-term and short-term…

计算机视觉与模式识别 · 计算机科学 2017-11-17 Tanya Marwah , Gaurav Mittal , Vineeth N. Balasubramanian
‹ 上一页 1 8 9 10 下一页 ›