中文
相关论文

相关论文: A Short Note about Kinetics-600

200 篇论文

Understanding the performance of machine learning models across diverse data distributions is critically important for reliable applications. Motivated by this, there is a growing focus on curating benchmark datasets that capture…

机器学习 · 计算机科学 2022-02-15 Weixin Liang , James Zou

This paper presents a framework to automate the labelling process for gestures in musical performance videos with a 3D Convolutional Neural Network (CNN). While this idea was proposed in a previous study, this paper introduces several…

计算机视觉与模式识别 · 计算机科学 2022-05-25 Foteini Simistira Liwicki , Richa Upadhyay , Prakash Chandra Chhipa , Killian Murphy , Federico Visi , Stefan Östersjö , Marcus Liwicki

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The…

计算机视觉与模式识别 · 计算机科学 2018-10-30 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

AI-generated video generation continues its journey through the uncanny valley to produce content that is increasingly perceptually indistinguishable from reality. To better protect individuals, organizations, and societies from its…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Matyas Bohacek , Hany Farid

In the field of robotic manipulation, deep imitation learning is recognized as a promising approach for acquiring manipulation skills. Additionally, learning from diverse robot datasets is considered a viable method to achieve versatility…

机器人学 · 计算机科学 2024-03-20 Heecheol Kim , Yoshiyuki Ohmura , Yasuo Kuniyoshi

Recent advancements in deep learning, computer vision, and embodied AI have given rise to synthetic causal reasoning video datasets. These datasets facilitate the development of AI algorithms that can reason about physical interactions…

人工智能 · 计算机科学 2021-08-16 Jiafei Duan , Samson Yu Bai Jian , Cheston Tan

The Meta Video Dataset (MetaVD) provides annotated relations between action classes in major datasets for human action recognition in videos. Although these annotated relations enable dataset augmentation, it is only applicable to those…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Yuya Yoshikawa , Yutaro Shigeto , Masashi Shimbo , Akikazu Takeuchi

Learning text-video embeddings usually requires a dataset of video clips with manually provided captions. However, such datasets are expensive and time consuming to create and therefore difficult to obtain on a large scale. In this work, we…

计算机视觉与模式识别 · 计算机科学 2019-08-01 Antoine Miech , Dimitri Zhukov , Jean-Baptiste Alayrac , Makarand Tapaswi , Ivan Laptev , Josef Sivic

What is the right way to reason about human activities? What directions forward are most promising? In this work, we analyze the current state of human activity understanding in videos. The goal of this paper is to examine datasets,…

计算机视觉与模式识别 · 计算机科学 2017-08-10 Gunnar A. Sigurdsson , Olga Russakovsky , Abhinav Gupta

The acquisition cost for large, annotated motion datasets remains a critical bottleneck for skeletal-based Human Activity Recognition (HAR). Although Text-to-Motion (T2M) generative models offer a compelling, scalable source of synthetic…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Luca Cazzola , Ahed Alboody

This technical report introduces our 2nd place solution to Kinetics-TPS Track on Part-level Action Parsing in ICCV DeeperAction Workshop 2021. Our entry is mainly based on YOLOF for instance and part detection, HRNet for human pose…

计算机视觉与模式识别 · 计算机科学 2022-09-02 Xiaodong Chen , Xinchen Liu , Kun Liu , Wu Liu , Tao Mei

Technologies play an increasingly important role in sports and become a real competitive advantage for the athletes who benefit from it. Among them, the use of motion capture is developing in various sports to optimize sporting gestures.…

计算机视觉与模式识别 · 计算机科学 2023-10-09 Fiche Guénolé , Sevestre Vincent , Gonzalez-Barral Camila , Leglaive Simon , Séguier Renaud

Advances in neural fields are enabling high-fidelity capture of the shape and appearance of dynamic 3D scenes. However, their capabilities lag behind those offered by conventional representations such as 2D videos because of algorithmic…

Driving datasets accelerate the development of intelligent driving and related computer vision technologies, while substantial and detailed annotations serve as fuels and powers to boost the efficacy of such datasets to improve…

机器学习 · 计算机科学 2019-06-04 Zhengping Che , Guangyu Li , Tracy Li , Bo Jiang , Xuefeng Shi , Xinsheng Zhang , Ying Lu , Guobin Wu , Yan Liu , Jieping Ye

We study the task of robust feature representations, aiming to generalize well on multiple datasets for action recognition. We build our method on Transformers for its efficacy. Although we have witnessed great progress for video action…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Junwei Liang , Enwei Zhang , Jun Zhang , Chunhua Shen

Facial expressions and actions differ among different individuals at varying degrees of intensity given responses to external stimuli, particularly among those that are neurodivergent. Such behaviors affect people in terms of overall…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Manuel Serna-Aguilera , Xuan Bac Nguyen , Han-Seok Seo , Khoa Luu

Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Yufan Deng , Daquan Zhou

This is the third year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human annotated training labels available for both passage and document ranking tasks. In…

信息检索 · 计算机科学 2025-07-14 Nick Craswell , Bhaskar Mitra , Emine Yilmaz , Daniel Campos , Jimmy Lin

Understanding human actions in visual data is tied to advances in complementary research areas including object recognition, human dynamics, domain adaptation and semantic segmentation. Over the last decade, human action analysis evolved…

计算机视觉与模式识别 · 计算机科学 2017-02-02 Samitha Herath , Mehrtash Harandi , Fatih Porikli

We introduce PointOdyssey, a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms. Our goal is to advance the state-of-the-art by placing emphasis on…

计算机视觉与模式识别 · 计算机科学 2023-07-28 Yang Zheng , Adam W. Harley , Bokui Shen , Gordon Wetzstein , Leonidas J. Guibas