中文
相关论文

相关论文: Hollywood in Homes: Crowdsourcing Data Collection …

200 篇论文

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset contains 8.5K…

计算机视觉与模式识别 · 计算机科学 2024-08-12 Tingkai Liu , Yunzhe Tao , Haogeng Liu , Qihang Fan , Ding Zhou , Huaibo Huang , Ran He , Hongxia Yang

Audio event detection is a widely studied audio processing task, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many efforts typically…

音频与语音处理 · 电气工程与系统科学 2023-02-16 Rajat Hebbar , Digbalay Bose , Krishna Somandepalli , Veena Vijai , Shrikanth Narayanan

Automatically identifying harmful content in video is an important task with a wide range of applications. However, there is a lack of professionally labeled open datasets available. In this work VidHarm, an open dataset of 3589 video clips…

计算机视觉与模式识别 · 计算机科学 2022-09-07 Johan Edstedt , Amanda Berg , Michael Felsberg , Johan Karlsson , Francisca Benavente , Anette Novak , Gustav Grund Pihlgren

Smart City applications such as intelligent traffic routing or accident prevention rely on computer vision methods for exact vehicle localization and tracking. Due to the scarcity of accurately labeled data, detecting and tracking vehicles…

计算机视觉与模式识别 · 计算机科学 2022-08-31 Fabian Herzog , Junpeng Chen , Torben Teepe , Johannes Gilg , Stefan Hörmann , Gerhard Rigoll

Motivated by the application of fact-level image understanding, we present an automatic method for data collection of structured visual facts from images with captions. Example structured facts include attributed objects (e.g., <flower,…

计算与语言 · 计算机科学 2016-04-11 Mohamed Elhoseiny , Scott Cohen , Walter Chang , Brian Price , Ahmed Elgammal

Large datasets of paired images and text have become increasingly popular for learning generic representations for vision and vision-and-language tasks. Such datasets have been built by querying search engines or collecting HTML alt-text --…

计算机视觉与模式识别 · 计算机科学 2021-11-23 Karan Desai , Gaurav Kaul , Zubin Aysola , Justin Johnson

Data scarcity presents a key bottleneck for imitation learning in robotic manipulation. In this paper, we focus on random exploration data-actions and video sequences produced autonomously via motions to randomly sampled positions in the…

机器人学 · 计算机科学 2025-09-19 Shutong Jin , Axel Kaliff , Ruiyu Wang , Muhammad Zahid , Florian T. Pokorny

We introduce UCF101 which is currently the largest dataset of human actions. It consists of 101 action classes, over 13k clips and 27 hours of video data. The database consists of realistic user uploaded videos containing camera motion and…

计算机视觉与模式识别 · 计算机科学 2012-12-04 Khurram Soomro , Amir Roshan Zamir , Mubarak Shah

Objective monitoring and assessment of human motor behavior can improve the diagnosis and management of several medical conditions. Over the past decade, significant advances have been made in the use of wearable technology for continuously…

计算机视觉与模式识别 · 计算机科学 2019-09-23 Behnaz Rezaei , Yiorgos Christakis , Bryan Ho , Kevin Thomas , Kelley Erb , Sarah Ostadabbas , Shyamal Patel

Many recent advancements in Computer Vision are attributed to large datasets. Open-source software packages for Machine Learning and inexpensive commodity hardware have reduced the barrier of entry for exploring novel approaches at scale.…

计算机视觉与模式识别 · 计算机科学 2016-09-29 Sami Abu-El-Haija , Nisarg Kothari , Joonseok Lee , Paul Natsev , George Toderici , Balakrishnan Varadarajan , Sudheendra Vijayanarasimhan

Diversity-aware data are essential for a robust modeling of human behavior in context. In addition, being the human behavior of interest for numerous applications, data must also be reusable across domain, to ensure diversity of…

计算机与社会 · 计算机科学 2023-06-19 Matteo Busso , Xiaoyue Li

Existing image/video datasets for cattle behavior recognition are mostly small, lack well-defined labels, or are collected in unrealistic controlled environments. This limits the utility of machine learning (ML) models learned from them.…

计算机视觉与模式识别 · 计算机科学 2023-07-04 Ali Zia , Renuka Sharma , Reza Arablouei , Greg Bishop-Hurley , Jody McNally , Neil Bagnall , Vivien Rolland , Brano Kusy , Lars Petersson , Aaron Ingham

Movie screenplay summarization is challenging, as it requires an understanding of long input contexts and various elements unique to movies. Large language models have shown significant advancements in document summarization, but they often…

计算与语言 · 计算机科学 2024-08-13 Rohit Saxena , Frank Keller

This paper introduces a new video-and-language dataset with human actions for multimodal logical inference, which focuses on intentional and aspectual expressions that describe dynamic human actions. The dataset consists of 200 videos,…

计算机视觉与模式识别 · 计算机科学 2021-06-29 Riko Suzuki , Hitomi Yanaka , Koji Mineshima , Daisuke Bekki

Knowing where people look is a useful tool in many various image and video applications. However, traditional gaze tracking hardware is expensive and requires local study participants, so acquiring gaze location data from a large number of…

社会与信息网络 · 计算机科学 2012-04-17 Dmitry Rudoy , Dan B. Goldman , Eli Shechtman , Lihi Zelnik-Manor

Understanding movies and their structural patterns is a crucial task in decoding the craft of video editing. While previous works have developed tools for general analysis, such as detecting characters or recognizing cinematography…

计算机视觉与模式识别 · 计算机科学 2022-10-25 Alejandro Pardo , Fabian Caba Heilbron , Juan León Alcázar , Ali Thabet , Bernard Ghanem

Most existing video tasks related to "human" focus on the segmentation of salient humans, ignoring the unspecified others in the video. Few studies have focused on segmenting and tracking all humans in a complex video, including pedestrians…

计算机视觉与模式识别 · 计算机科学 2021-08-17 Ran Yu , Chenyu Tian , Weihao Xia , Xinyuan Zhao , Haoqian Wang , Yujiu Yang

This paper addresses automatic summarization and search in visual data comprising of videos, live streams and image collections in a unified manner. In particular, we propose a framework for multi-faceted summarization which extracts…

计算机视觉与模式识别 · 计算机科学 2017-04-06 Anurag Sahoo , Vishal Kaushal , Khoshrav Doctor , Suyash Shetty , Rishabh Iyer , Ganesh Ramakrishnan

Most existing robotic datasets capture static scene data and thus are limited in evaluating robots' dynamic performance. To address this, we present a mobile robot oriented large-scale indoor dataset, denoted as THUD (Tsinghua University…

机器人学 · 计算机科学 2024-07-02 Yifan Tang , Cong Tai , Fangxing Chen , Wanting Zhang , Tao Zhang , Xueping Liu , Yongjin Liu , Long Zeng

We introduce DAiSEE, the first multi-label video classification dataset comprising of 9068 video snippets captured from 112 users for recognizing the user affective states of boredom, confusion, engagement, and frustration in the wild. The…

计算机视觉与模式识别 · 计算机科学 2022-07-08 Abhay Gupta , Arjun D'Cunha , Kamal Awasthi , Vineeth Balasubramanian