中文
相关论文

相关论文: Video Caption Dataset for Describing Human Actions…

200 篇论文

Tutorial videos of mobile apps have become a popular and compelling way for users to learn unfamiliar app features. To make the video accessible to the users, video creators always need to annotate the actions in the video, including what…

人机交互 · 计算机科学 2023-08-08 Sidong Feng , Chunyang Chen , Zhenchang Xing

Cognitive science has shown that humans perceive videos in terms of events separated by the state changes of dominant subjects. State changes trigger new events and are one of the most useful among the large amount of redundant information…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Yuxuan Wang , Difei Gao , Licheng Yu , Stan Weixian Lei , Matt Feiszli , Mike Zheng Shou

Generating high-quality cartoon animations multimodal control is challenging due to the complexity of non-human characters, stylistically diverse motions and fine-grained emotions. There is a huge domain gap between real-world videos and…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Shuolin Xu , Bingyuan Wang , Zeyu Cai , Fangteng Fu , Yue Ma , Tongyi Lee , Hongchuan Yu , Zeyu Wang

Lifelogging cameras capture everyday life from a first-person perspective, but generate so much data that it is hard for users to browse and organize their image collections effectively. In this paper, we propose to use automatic image…

计算机视觉与模式识别 · 计算机科学 2016-08-15 Chenyou Fan , David J. Crandall

Some datasets with the described content and order of occurrence of sounds have been released for conversion between environmental sound and text. However, there are very few texts that include information on the impressions humans feel,…

We present an efficient framework that can generate a coherent paragraph to describe a given video. Previous works on video captioning usually focus on video clips. They typically treat an entire video as a whole and generate the caption…

计算机视觉与模式识别 · 计算机科学 2018-07-27 Yilei Xiong , Bo Dai , Dahua Lin

A great video title describes the most salient event compactly and captures the viewer's attention. In contrast, video captioning tends to generate sentences that describe the video as a whole. Although generating a video title…

计算机视觉与模式识别 · 计算机科学 2016-09-09 Kuo-Hao Zeng , Tseng-Hung Chen , Juan Carlos Niebles , Min Sun

Accelerated by the tremendous increase in Internet bandwidth and storage space, video data has been generated, published and spread explosively, becoming an indispensable part of today's big data. In this paper, we focus on reviewing two…

计算机视觉与模式识别 · 计算机科学 2018-02-23 Zuxuan Wu , Ting Yao , Yanwei Fu , Yu-Gang Jiang

Deep video action recognition models have been highly successful in recent years but require large quantities of manually annotated data, which are expensive and laborious to obtain. In this work, we investigate the generation of synthetic…

计算机视觉与模式识别 · 计算机科学 2019-10-16 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Naila Murray , Antonio Manuel López

This paper discusses and demonstrates the outcomes from our experimentation on Image Captioning. Image captioning is a much more involved task than image recognition or classification, because of the additional challenge of recognizing the…

计算机视觉与模式识别 · 计算机科学 2018-05-24 Vikram Mullachery , Vishal Motwani

Video captioning is an essential technology to understand scenes and describe events in natural language. To apply it to real-time monitoring, a system needs not only to describe events accurately but also to produce the captions as soon as…

计算机视觉与模式识别 · 计算机科学 2021-08-05 Chiori Hori , Takaaki Hori , Jonathan Le Roux

Generating a description of an image is called image captioning. Image captioning requires to recognize the important objects, their attributes and their relationships in an image. It also needs to generate syntactically and semantically…

计算机视觉与模式识别 · 计算机科学 2018-10-16 Md. Zakir Hossain , Ferdous Sohel , Mohd Fairuz Shiratuddin , Hamid Laga

In this work, we present a novel dataset consisting of eye movements and verbal descriptions recorded synchronously over images. Using this data, we study the differences in human attention during free-viewing and image captioning tasks. We…

计算机视觉与模式识别 · 计算机科学 2019-08-08 Sen He , Hamed R. Tavakoli , Ali Borji , Nicolas Pugeault

We propose a method for human action recognition, one that can localize the spatiotemporal regions that `define' the actions. This is a challenging task due to the subtlety of human actions in video and the co-occurrence of contextual…

计算机视觉与模式识别 · 计算机科学 2019-04-12 Yang Wang , Vinh Tran , Gedas Bertasius , Lorenzo Torresani , Minh Hoai

Counting repetitive actions are widely seen in human activities such as physical exercise. Existing methods focus on performing repetitive action counting in short videos, which is tough for dealing with longer videos in more realistic…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Huazhang Hu , Sixun Dong , Yiqun Zhao , Dongze Lian , Zhengxin Li , Shenghua Gao

Most of the existing works on human activity analysis focus on recognition or early recognition of the activity labels from complete or partial observations. Similarly, almost all of the existing video captioning approaches focus on the…

计算机视觉与模式识别 · 计算机科学 2021-05-28 Tahmida Mahmud , Mohammad Billah , Mahmudul Hasan , Amit K. Roy-Chowdhury

Current deep learning models often achieve excellent results on benchmark image-to-text datasets but fail to generate texts that are useful in practice. We argue that to close this gap, it is vital to distinguish descriptions from captions…

计算与语言 · 计算机科学 2022-10-31 Elisa Kreiss , Fei Fang , Noah D. Goodman , Christopher Potts

Body and face motion play an integral role in communication. They convey crucial information on the participants. Advances in generative modeling and multi-modal learning have enabled motion generation from signals such as speech,…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Lownish Rai Sookha , Nikhil Pakhale , Mudasir Ganaie , Abhinav Dhall

We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally consistent bounding boxes.…

计算机视觉与模式识别 · 计算机科学 2024-11-13 Evangelos Kazakos , Cordelia Schmid , Josef Sivic

Creating and labelling datasets of videos for use in training Human Activity Recognition models is an arduous task. In this paper, we approach this by using 3D rendering tools to generate a synthetic dataset of videos, and show that a…

计算机视觉与模式识别 · 计算机科学 2020-07-23 Ollie Matthews , Koki Ryu , Tarun Srivastava