English
Related papers

Related papers: M-VAD Names: a Dataset for Video Captioning with N…

200 papers

The Audio Description (AD) task aims to generate descriptions of visual elements for visually impaired individuals to help them access long-form video content, like movies. With video feature, text, character bank and context information as…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Hanlin Wang , Zhan Tong , Kecheng Zheng , Yujun Shen , Limin Wang

Video description involves the generation of the natural language description of actions, events, and objects in the video. There are various applications of video description by filling the gap between languages and vision for visually…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Alok Singh , Thoudam Doren Singh , Sivaji Bandyopadhyay

An ideal description for a given video should fix its gaze on salient and representative content, which is capable of distinguishing this video from others. However, the distribution of different words is unbalanced in video captioning…

Computer Vision and Pattern Recognition · Computer Science 2019-01-03 Jiarong Dong , Ke Gao , Xiaokai Chen , Junbo Guo , Juan Cao , Yongdong Zhang

Image captioning is a multimodal problem that has drawn extensive attention in both the natural language processing and computer vision community. In this paper, we present a novel image captioning architecture to better explore semantics…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Zhan Shi , Xu Zhou , Xipeng Qiu , Xiaodan Zhu

Soft-biometrics play an important role in face biometrics and related fields since these might lead to biased performances, threatens the user's privacy, or are valuable for commercial aspects. Current face databases are specifically…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Philipp Terhörst , Daniel Fährmann , Jan Niklas Kolf , Naser Damer , Florian Kirchbuchner , Arjan Kuijper

This paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. The dataset consists of images selected to unambiguously illustrate…

Computation and Language · Computer Science 2022-06-20 Josiah Wang , Pranava Madhyastha , Josiel Figueiredo , Chiraag Lala , Lucia Specia

True understanding of videos comes from a joint analysis of all its modalities: the video frames, the audio track, and any accompanying text such as closed captions. We present a way to learn a compact multimodal feature representation that…

Computer Vision and Pattern Recognition · Computer Science 2020-04-07 Vivek Sharma , Makarand Tapaswi , Rainer Stiefelhagen

Recent progress in face detection (including keypoint detection), and recognition is mainly being driven by (i) deeper convolutional neural network architectures, and (ii) larger datasets. However, most of the large datasets are maintained…

Computer Vision and Pattern Recognition · Computer Science 2017-05-23 Ankan Bansal , Anirudh Nanduri , Carlos Castillo , Rajeev Ranjan , Rama Chellappa

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Kun Li , Pengyu Liu , Dan Guo , Fei Wang , Zhiliang Wu , Hehe Fan , Meng Wang

The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected with overrestrictive…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Soravit Changpinyo , Piyush Sharma , Nan Ding , Radu Soricut

Large-scale annotated datasets allow AI systems to learn from and build upon the knowledge of the crowd. Many crowdsourcing techniques have been developed for collecting image annotations. These techniques often implicitly rely on the fact…

Human-Computer Interaction · Computer Science 2016-10-07 Gunnar A. Sigurdsson , Olga Russakovsky , Ali Farhadi , Ivan Laptev , Abhinav Gupta

This paper presents a new video question answering task on screencast tutorials. We introduce a dataset including question, answer and context triples from the tutorial videos for a software. Unlike other video question answering works, all…

Computation and Language · Computer Science 2020-08-04 Wentian Zhao , Seokhwan Kim , Ning Xu , Hailin Jin

Robots are becoming everyday devices, increasing their interaction with humans. To make human-machine interaction more natural, cognitive features like Visual Voice Activity Detection (VVAD), which can detect whether a person is speaking or…

Computer Vision and Pattern Recognition · Computer Science 2024-01-03 Adrian Lubitz , Matias Valdenegro-Toro , Frank Kirchner

Learning a joint language-visual embedding has a number of very appealing properties and can result in variety of practical application, including natural language image/video annotation and search. In this work, we study three different…

Computer Vision and Pattern Recognition · Computer Science 2016-09-27 Atousa Torabi , Niket Tandon , Leonid Sigal

Although the performance of person Re-Identification (ReID) has been significantly boosted, many challenging issues in real scenarios have not been fully investigated, e.g., the complex scenes and lighting variations, viewpoint and pose…

Computer Vision and Pattern Recognition · Computer Science 2018-06-26 Longhui Wei , Shiliang Zhang , Wen Gao , Qi Tian

Humans share a strong tendency to memorize/forget some of the visual information they encounter. This paper focuses on providing computational models for the prediction of the intrinsic memorability of visual content. To address this new…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Romain Cohendet , Claire-Hélène Demarty , Ngoc Q. K. Duong , Martin Engilberge

Given the features of a video, recurrent neural networks can be used to automatically generate a caption for the video. Existing methods for video captioning have at least three limitations. First, semantic information has been widely…

Computer Vision and Pattern Recognition · Computer Science 2021-02-15 Haoran Chen , Ke Lin , Alexander Maye , Jianming Li , Xiaolin Hu

Visual Object Tracking (VOT) is a fundamental task with widespread applications in autonomous navigation, surveillance, and maritime robotics. Despite significant advances in generic object tracking, maritime environments continue to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Ahsan Baidar Bakht , Muhayy Ud Din , Sajid Javed , Irfan Hussain

Annotation of multimedia data by humans is time-consuming and costly, while reliable automatic generation of semantic metadata is a major challenge. We propose a framework to extract semantic metadata from automatically generated video…

Computer Vision and Pattern Recognition · Computer Science 2023-09-14 Johannes Scherer , Ansgar Scherp , Deepayan Bhowmik

Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both…

Computer Vision and Pattern Recognition · Computer Science 2017-05-03 Ranjay Krishna , Kenji Hata , Frederic Ren , Li Fei-Fei , Juan Carlos Niebles