English
Related papers

Related papers: Hollywood in Homes: Crowdsourcing Data Collection …

200 papers

We introduce QuerYD, a new large-scale dataset for retrieval and event localisation in video. A unique feature of our dataset is the availability of two audio tracks for each video: the original audio, and a high-quality spoken description…

Computer Vision and Pattern Recognition · Computer Science 2021-02-18 Andreea-Maria Oncescu , João F. Henriques , Yang Liu , Andrew Zisserman , Samuel Albanie

Art forms such as movies and television (TV) dramas are reflections of the real world, which have attracted much attention from the multimodal learning community recently. However, existing corpora in this domain share three limitations:…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Chen Li , Xutan Peng , Teng Wang , Yixiao Ge , Mengyang Liu , Xuyuan Xu , Yexin Wang , Ying Shan

We describe the 2020 edition of the DeepMind Kinetics human action dataset, which replenishes and extends the Kinetics-700 dataset. In this new version, there are at least 700 video clips from different YouTube videos for each of the 700…

Computer Vision and Pattern Recognition · Computer Science 2020-10-22 Lucas Smaira , João Carreira , Eric Noland , Ellen Clancy , Amy Wu , Andrew Zisserman

We address the problem of accurate capture of interactive behaviors between two people in daily scenarios. Most previous works either only consider one person or solely focus on conversational gestures of two people, assuming the body…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Leo Ho , Yinghao Huang , Dafei Qin , Mingyi Shi , Wangpok Tse , Wei Liu , Junichi Yamagishi , Taku Komura

Humans have long been recorded in a variety of forms since antiquity. For example, sculptures and paintings were the primary media for depicting human beings before the invention of cameras. However, most current human-centric computer…

Computer Vision and Pattern Recognition · Computer Science 2023-04-06 Xuan Ju , Ailing Zeng , Jianan Wang , Qiang Xu , Lei Zhang

Crowdsourcing has been widely used to efficiently obtain labeled datasets for supervised learning from large numbers of human resources at low cost. However, one of the technical challenges in obtaining high-quality results from…

Human-Computer Interaction · Computer Science 2023-02-28 Ryosuke Ueda , Koh Takeuchi , Hisashi Kashima

To enable machines to understand the way humans interact with the physical world in daily life, 3D interaction signals should be captured in natural settings, allowing people to engage with multiple objects in a range of sequential and…

Computer Vision and Pattern Recognition · Computer Science 2025-01-23 Jeonghwan Kim , Jisoo Kim , Jeonghyeon Na , Hanbyul Joo

Despite recent advances of AI, story understanding remains an open and under-investigated problem. We collect, preprocess, and publicly release a video-language story dataset, Synopses of Movie Narratives (SyMoN), containing 5,193 video…

Computer Vision and Pattern Recognition · Computer Science 2023-04-06 Yidan Sun , Qin Chao , Yangfeng Ji , Boyang Li

Event perception tasks such as recognizing and localizing actions in streaming videos are essential for scaling to real-world application contexts. We tackle the problem of learning actor-centered representations through the notion of…

Computer Vision and Pattern Recognition · Computer Science 2022-12-01 Sathyanarayanan N. Aakur , Sudeep Sarkar

Understanding the emotional impact of movies has become important for affective movie analysis, ranking, and indexing. Methods for recognizing evoked emotions are usually trained on human annotated data. Concretely, viewers watch video…

Computer Vision and Pattern Recognition · Computer Science 2021-08-02 Hassan Hayat , Carles Ventura , Agata Lapedriza

We consider the task of identifying human actions visible in online videos. We focus on the widely spread genre of lifestyle vlogs, which consist of videos of people performing actions while verbally describing them. Our goal is to identify…

Computation and Language · Computer Science 2021-09-10 Oana Ignat , Laura Burdick , Jia Deng , Rada Mihalcea

This paper presents two novel approaches for people counting in crowded and open environments that combine the information gathered by multiple views. Multiple camera are used to expand the field of view as well as to mitigate the problem…

Computer Vision and Pattern Recognition · Computer Science 2017-05-09 Fabio Dittrich , Luiz E. S. de Oliveira , Alceu S. Britto , Alessandro L. Koerich

Robust face clustering is a vital step in enabling computational understanding of visual character portrayal in media. Face clustering for long-form content is challenging because of variations in appearance and lack of supporting…

Computer Vision and Pattern Recognition · Computer Science 2022-03-01 Krishna Somandepalli , Rajat Hebbar , Shrikanth Narayanan

This paper focuses on the temporal aspect for recognizing human activities in videos; an important visual cue that has long been undervalued. We revisit the conventional definition of activity and restrict it to Complex Action: a set of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-30 Noureldien Hussein , Efstratios Gavves , Arnold W. M. Smeulders

Eye-in-hand cameras have shown promise in enabling greater sample efficiency and generalization in vision-based robotic manipulation. However, for robotic imitation, it is still expensive to have a human teleoperator collect large amounts…

Robotics · Computer Science 2023-07-13 Moo Jin Kim , Jiajun Wu , Chelsea Finn

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully,…

Computer Vision and Pattern Recognition · Computer Science 2019-05-10 Huda Alamri , Vincent Cartillier , Abhishek Das , Jue Wang , Anoop Cherian , Irfan Essa , Dhruv Batra , Tim K. Marks , Chiori Hori , Peter Anderson , Stefan Lee , Devi Parikh

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Jiaben Chen , Zixin Wang , Ailing Zeng , Yang Fu , Xueyang Yu , Siyuan Cen , Julian Tanke , Yihang Chen , Koichi Saito , Yuki Mitsufuji , Chuang Gan

The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Mattia Soldan , Alejandro Pardo , Juan León Alcázar , Fabian Caba Heilbron , Chen Zhao , Silvio Giancola , Bernard Ghanem

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to…

Computation and Language · Computer Science 2023-09-26 Michele Cafagna , Kees van Deemter , Albert Gatt

3D scene understanding is a long-standing challenge in computer vision and a key component in enabling mixed reality, wearable computing, and embodied AI. Providing a solution to these applications requires a multifaceted approach that…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Anna-Maria Halacheva , Yang Miao , Jan-Nico Zaech , Xi Wang , Luc Van Gool , Danda Pani Paudel
‹ Prev 1 8 9 10 Next ›