English
Related papers

Related papers: Building a Video-and-Language Dataset with Human A…

200 papers

We introduce a new task, Video-and-Language Inference, for joint multimodal understanding of video and text. Given a video clip with aligned subtitles as premise, paired with a natural language hypothesis based on the video content, a model…

Computer Vision and Pattern Recognition · Computer Science 2020-03-27 Jingzhou Liu , Wenhu Chen , Yu Cheng , Zhe Gan , Licheng Yu , Yiming Yang , Jingjing Liu

We aim to automatically identify human action reasons in online videos. We focus on the widespread genre of lifestyle vlogs, in which people perform actions while verbally describing them. We introduce and make publicly available the WhyAct…

Computer Vision and Pattern Recognition · Computer Science 2021-09-10 Oana Ignat , Santiago Castro , Hanwen Miao , Weiji Li , Rada Mihalcea

In recent years, automatic video caption generation has attracted considerable attention. This paper focuses on the generation of Japanese captions for describing human actions. While most currently available video caption datasets have…

Computation and Language · Computer Science 2020-03-11 Yutaro Shigeto , Yuya Yoshikawa , Jiaqing Lin , Akikazu Takeuchi

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to…

Computation and Language · Computer Science 2023-09-26 Michele Cafagna , Kees van Deemter , Albert Gatt

We describe the DeepMind Kinetics human action video dataset. The dataset contains 400 human action classes, with at least 400 video clips for each action. Each clip lasts around 10s and is taken from a different YouTube video. The actions…

This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Baoyao Yang , Wanyun Li , Dixin Chen , Junxiang Chen , Wenbin Yao , Haifeng Lin

Although empathic interaction between counselor and client is fundamental to success in the psychotherapeutic process, there are currently few datasets to aid a computational approach to empathy understanding. In this paper, we construct a…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Zhou'an_Zhu , Xin Li , Jicai Pan , Yufei Xiao , Yanan Chang , Feiyi Zheng , Shangfei Wang

Videos capture events that typically contain multiple sequential, and simultaneous, actions even in the span of only a few seconds. However, most large-scale datasets built to train models for action recognition in video only provide a…

Computer Vision and Pattern Recognition · Computer Science 2021-09-29 Mathew Monfort , Bowen Pan , Kandan Ramakrishnan , Alex Andonian , Barry A McNamara , Alex Lascelles , Quanfu Fan , Dan Gutfreund , Rogerio Feris , Aude Oliva

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions…

Computer Vision and Pattern Recognition · Computer Science 2019-02-19 Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan , Lisa Brown , Quanfu Fan , Dan Gutfruend , Carl Vondrick , Aude Oliva

What is the right way to reason about human activities? What directions forward are most promising? In this work, we analyze the current state of human activity understanding in videos. The goal of this paper is to examine datasets,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-10 Gunnar A. Sigurdsson , Olga Russakovsky , Abhinav Gupta

Deep learning for human action recognition in videos is making significant progress, but is slowed down by its dependency on expensive manual labeling of large video collections. In this work, we investigate the generation of synthetic…

Computer Vision and Pattern Recognition · Computer Science 2017-07-20 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Antonio Manuel López Peña

Advancements in deep neural networks have contributed to near perfect results for many computer vision problems such as object recognition, face recognition and pose estimation. However, human action recognition is still far from…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Asanka G. Perera , Yee Wei Law , Titilayo T. Ogunwa , Javaan Chahl

Text-level discourse parsing aims to unmask how two sentences in the text are related to each other. We propose the task of Visual Discourse Parsing, which requires understanding discourse relations among scenes in a video. Here we use the…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Arjun R. Akula , Song-Chun Zhu

In the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because describing each…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Lingru Zhou , Yiqi Gao , Manqing Zhang , Peng Wu , Peng Wang , Yanning Zhang

This paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. The dataset consists of images selected to unambiguously illustrate…

Computation and Language · Computer Science 2022-06-20 Josiah Wang , Pranava Madhyastha , Josiel Figueiredo , Chiraag Lala , Lucia Specia

We address the problem of accurate capture of interactive behaviors between two people in daily scenarios. Most previous works either only consider one person or solely focus on conversational gestures of two people, assuming the body…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Leo Ho , Yinghao Huang , Dafei Qin , Mingyi Shi , Wangpok Tse , Wei Liu , Junichi Yamagishi , Taku Komura

We consider the task of identifying human actions visible in online videos. We focus on the widely spread genre of lifestyle vlogs, which consist of videos of people performing actions while verbally describing them. Our goal is to identify…

Computation and Language · Computer Science 2021-09-10 Oana Ignat , Laura Burdick , Jia Deng , Rada Mihalcea

Understanding human intentions is key to enabling effective and efficient human-robot interaction (HRI) in collaborative settings. To enable developments and evaluation of the ability of artificial intelligence (AI) systems to infer human…

Computer Vision and Pattern Recognition · Computer Science 2023-05-01 Jiafei Duan , Samson Yu , Nicholas Tan , Yi Ru Wang , Cheston Tan

Human Action Recognition (HAR) is a very crucial task in computer vision. It helps to carry out a series of downstream tasks, like understanding human behaviors. Due to the complexity of human behaviors, many highly valuable behaviors are…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Hongwu Li , Zhenliang Zhang , Wei Wang

Captioning is a crucial and challenging task for video understanding. In videos that involve active agents such as humans, the agent's actions can bring about myriad changes in the scene. Observable changes such as movements, manipulations,…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Zhiyuan Fang , Tejas Gokhale , Pratyay Banerjee , Chitta Baral , Yezhou Yang
‹ Prev 1 2 3 10 Next ›