中文
相关论文

相关论文: A Short Note about Kinetics-600

200 篇论文

We introduce DynaSent ('Dynamic Sentiment'), a new English-language benchmark task for ternary (positive/negative/neutral) sentiment analysis. DynaSent combines naturally occurring sentences with sentences created using the open-source…

计算与语言 · 计算机科学 2021-01-01 Christopher Potts , Zhengxuan Wu , Atticus Geiger , Douwe Kiela

Modeling crowd behavior relies on accurate data of pedestrian movements at a high level of detail. Imaging sensors such as cameras provide a good basis for capturing such detailed pedestrian motion data. However, currently available…

计算机视觉与模式识别 · 计算机科学 2012-10-11 Stefan Seer , Norbert Brändle , Carlo Ratti

Understanding bimanual human hand activities is a critical problem in AI and robotics. We cannot build large models of bimanual activities because existing datasets lack the scale, coverage of diverse hand activities, and detailed…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Rao Fu , Dingxi Zhang , Alex Jiang , Wanjia Fu , Austin Funk , Daniel Ritchie , Srinath Sridhar

Despite the fact that many 3D human activity benchmarks being proposed, most existing action datasets focus on the action recognition tasks for the segmented videos. There is a lack of standard large-scale benchmarks, especially for current…

计算机视觉与模式识别 · 计算机科学 2017-03-29 Chunhui Liu , Yueyu Hu , Yanghao Li , Sijie Song , Jiaying Liu

We present a new publicly available dataset with the goal of advancing multi-modality learning by offering vision and language data within the same context. This is achieved by obtaining data from a social media website with posts…

计算与语言 · 计算机科学 2020-06-16 Bofan Xue , David Chan , John Canny

Classifying player actions from soccer videos is a challenging problem, which has become increasingly important in sports analytics over the years. Most state-of-the-art methods employ highly complex offline networks, which makes it…

计算机视觉与模式识别 · 计算机科学 2023-07-26 Sarosij Bose , Saikat Sarkar , Amlan Chakrabarti

The paper provides a survey of the development of machine-learning techniques for video analysis. The survey provides a summary of the most popular deep learning methods used for human activity recognition. We discuss how popular…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Marios S. Pattichis , Venkatesh Jatla , Alvaro E. Ullao Cerna

Understanding human visual attention and saliency is an integral part of vision research. In this context, there is an ever-present need for fresh and diverse benchmark datasets, particularly for insight into special use cases like crowded…

计算机视觉与模式识别 · 计算机科学 2019-10-10 Memoona Tahira , Sobas Mehboob , Anis U. Rahman , Omar Arif

Video action recognition is one of the representative tasks for video understanding. Over the last decade, we have witnessed great advancements in video action recognition thanks to the emergence of deep learning. But we also encountered…

计算机视觉与模式识别 · 计算机科学 2020-12-14 Yi Zhu , Xinyu Li , Chunhui Liu , Mohammadreza Zolfaghari , Yuanjun Xiong , Chongruo Wu , Zhi Zhang , Joseph Tighe , R. Manmatha , Mu Li

Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations…

计算机视觉与模式识别 · 计算机科学 2022-05-03 Fadime Sener , Dibyadip Chatterjee , Daniel Shelepov , Kun He , Dipika Singhania , Robert Wang , Angela Yao

Action recognition is so far mainly focusing on the problem of classification of hand selected preclipped actions and reaching impressive results in this field. But with the performance even ceiling on current datasets, it also appears that…

计算机视觉与模式识别 · 计算机科学 2019-06-05 Hilde Kuehne , Ahsan Iqbal , Alexander Richard , Juergen Gall

We find that existing language modeling datasets contain many near-duplicate examples and long repetitive substrings. As a result, over 1% of the unprompted output of language models trained on these datasets is copied verbatim from the…

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

计算机视觉与模式识别 · 计算机科学 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations. However, generating accurate lip-syncs degrades when applied to input speech in other languages, possibly due to the lack…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Kim Sung-Bin , Lee Chae-Yeon , Gihun Son , Oh Hyun-Bin , Janghoon Ju , Suekyeong Nam , Tae-Hyun Oh

Deep learning has shown remarkable progress in a wide range of problems. However, efficient training of such models requires large-scale datasets, and getting annotations for such datasets can be challenging and costly. In this work, we…

多媒体 · 计算机科学 2021-10-14 Mohit Sharma , Raj Patra , Harshal Desai , Shruti Vyas , Yogesh Rawat , Rajiv Ratn Shah

Human actions in videos are 3D signals. However, there are a few methods available for multiple human action recognition. For long videos, it's difficult to search within a video for a specific action and/or person. For that, this paper…

计算机视觉与模式识别 · 计算机科学 2021-03-16 Noor Almaadeed , Omar Elharrouss , Somaya Al-Maadeed , Ahmed Bouridane , Azeddine Beghdadi

Videos capture events that typically contain multiple sequential, and simultaneous, actions even in the span of only a few seconds. However, most large-scale datasets built to train models for action recognition in video only provide a…

计算机视觉与模式识别 · 计算机科学 2021-09-29 Mathew Monfort , Bowen Pan , Kandan Ramakrishnan , Alex Andonian , Barry A McNamara , Alex Lascelles , Quanfu Fan , Dan Gutfreund , Rogerio Feris , Aude Oliva

While reconstructing human poses in 3D from inexpensive sensors has advanced significantly in recent years, quantifying the dynamics of human motion, including the muscle-generated joint torques and external forces, remains a challenge.…

In this paper, we study current and upcoming frontiers across the landscape of skeleton-based human action recognition. To study skeleton-action recognition in the wild, we introduce Skeletics-152, a curated and 3-D pose-annotated subset of…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Pranay Gupta , Anirudh Thatipelli , Aditya Aggarwal , Shubh Maheshwari , Neel Trivedi , Sourav Das , Ravi Kiran Sarvadevabhatla

In this work, we propose Knowledge Integration Networks (referred as KINet) for video action recognition. KINet is capable of aggregating meaningful context features which are of great importance to identifying an action, such as human…

计算机视觉与模式识别 · 计算机科学 2020-02-19 Shiwen Zhang , Sheng Guo , Limin Wang , Weilin Huang , Matthew R. Scott