English
Related papers

Related papers: Video Caption Dataset for Describing Human Actions…

200 papers

Watching instructional videos are often used to learn about procedures. Video captioning is one way of automatically collecting such knowledge. However, it provides only an indirect, overall evaluation of multimodal models with no…

Computation and Language · Computer Science 2020-10-12 Frank F. Xu , Lei Ji , Botian Shi , Junyi Du , Graham Neubig , Yonatan Bisk , Nan Duan

We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2 million captions, collected through crowdsourcing rating data for The New Yorker's weekly cartoon caption…

Human action recognition and analysis have great demand and important application significance in video surveillance, video retrieval, and human-computer interaction. The task of human action quality evaluation requires the intelligent…

Computer Vision and Pattern Recognition · Computer Science 2022-04-21 Shunli Wang , Dingkang Yang , Peng Zhai , Qing Yu , Tao Suo , Zhan Sun , Ka Li , Lihua Zhang

In grammatical error correction (GEC), automatic evaluation is an important factor for research and development of GEC systems. Previous studies on automatic evaluation have demonstrated that quality estimation models built from datasets…

Computation and Language · Computer Science 2022-01-21 Daisuke Suzuki , Yujin Takahashi , Ikumi Yamashita , Taichi Aida , Tosho Hirasawa , Michitaka Nakatsuji , Masato Mita , Mamoru Komachi

An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos, predict rich, detailed textual descriptions, and be able to produce outputs before processing…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Xingyi Zhou , Anurag Arnab , Shyamal Buch , Shen Yan , Austin Myers , Xuehan Xiong , Arsha Nagrani , Cordelia Schmid

Derived from rapid advances in computer vision and machine learning, video analysis tasks have been moving from inferring the present state to predicting the future state. Vision-based action recognition and prediction from videos are such…

Computer Vision and Pattern Recognition · Computer Science 2022-02-15 Yu Kong , Yun Fu

While an important problem in the vision community is to design algorithms that can automatically caption images, few publicly-available datasets for algorithm development directly address the interests of real users. Observing that people…

Computer Vision and Pattern Recognition · Computer Science 2020-07-16 Danna Gurari , Yinan Zhao , Meng Zhang , Nilavra Bhattacharya

In this paper, we propose the task of live comment generation. Live comments are a new form of comments on videos, which can be regarded as a mixture of comments and chats. A high-quality live comment should be not only relevant to the…

Computation and Language · Computer Science 2018-08-14 Damai Dai

We examine the possibility that recent promising results in automatic caption generation are due primarily to language models. By varying image representation quality produced by a convolutional neural network, we find that a…

Computation and Language · Computer Science 2015-08-11 Jack Hessel , Nicolas Savva , Michael J. Wilber

We propose and demonstrate the task of giving natural language summaries of the actions of a robotic agent in a virtual environment. We explain why such a task is important, what makes it difficult, and discuss how it might be addressed. To…

Computation and Language · Computer Science 2022-03-15 Chad DeChant , Daniel Bauer

Automated audio captioning is a cross-modal translation task that aims to generate natural language descriptions for given audio clips. This task has received increasing attention with the release of freely available datasets in recent…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-28 Xinhao Mei , Xubo Liu , Mark D. Plumbley , Wenwu Wang

Videos capture events that typically contain multiple sequential, and simultaneous, actions even in the span of only a few seconds. However, most large-scale datasets built to train models for action recognition in video only provide a…

Computer Vision and Pattern Recognition · Computer Science 2021-09-29 Mathew Monfort , Bowen Pan , Kandan Ramakrishnan , Alex Andonian , Barry A McNamara , Alex Lascelles , Quanfu Fan , Dan Gutfreund , Rogerio Feris , Aude Oliva

There are substantial instructional videos on the Internet, which provide us tutorials for completing various tasks. Existing instructional video datasets only focus on specific steps at the video level, lacking experiential guidelines at…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Jiafeng Liang , Shixin Jiang , Zekun Wang , Haojie Pan , Zerui Chen , Zheng Chu , Ming Liu , Ruiji Fu , Zhongyuan Wang , Bing Qin

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements for high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Zhucun Xue , Jiangning Zhang , Teng Hu , Haoyang He , Yinan Chen , Yuxuan Cai , Yabiao Wang , Chengjie Wang , Yong Liu , Xiangtai Li , Dacheng Tao

With recent advances in computer vision and graphics, it is now possible to generate videos with extremely realistic synthetic faces, even in real time. Countless applications are possible, some of which raise a legitimate alarm, calling…

Computer Vision and Pattern Recognition · Computer Science 2018-03-28 Andreas Rössler , Davide Cozzolino , Luisa Verdoliva , Christian Riess , Justus Thies , Matthias Nießner

Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have…

Sound · Computer Science 2024-10-17 Mithun Manivannan , Vignesh Nethrapalli , Mark Cartwright

Understanding human actions is a key problem in computer vision. However, recognizing actions is only the first step of understanding what a person is doing. In this paper, we introduce the problem of predicting why a person has performed…

Computer Vision and Pattern Recognition · Computer Science 2016-12-01 Carl Vondrick , Deniz Oktay , Hamed Pirsiavash , Antonio Torralba

A large amount of recent research has focused on tasks that combine language and vision, resulting in a proliferation of datasets and methods. One such task is action recognition, whose applications include image annotation, scene under-…

Computation and Language · Computer Science 2017-04-25 Spandana Gella , Frank Keller

Generating video descriptions in natural language (a.k.a. video captioning) is a more challenging task than image captioning as the videos are intrinsically more complicated than images in two aspects. First, videos cover a broader range of…

Computer Vision and Pattern Recognition · Computer Science 2017-09-05 Shizhe Chen , Jia Chen , Qin Jin

Current deep learning results on video generation are limited while there are only a few first results on video prediction and no relevant significant results on video completion. This is due to the severe ill-posedness inherent in these…

Computer Vision and Pattern Recognition · Computer Science 2018-12-24 Haoye Cai , Chunyan Bai , Yu-Wing Tai , Chi-Keung Tang
‹ Prev 1 8 9 10 Next ›