中文
相关论文

相关论文: Action100M: A Large-scale Video Action Dataset

200 篇论文

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual…

计算机视觉与模式识别 · 计算机科学 2022-05-12 Antoine Yang , Antoine Miech , Josef Sivic , Ivan Laptev , Cordelia Schmid

Thanks to the substantial and explosively inscreased instructional videos on the Internet, novices are able to acquire knowledge for completing various tasks. Over the past decade, growing efforts have been devoted to investigating the…

计算机视觉与模式识别 · 计算机科学 2020-03-23 Yansong Tang , Jiwen Lu , Jie Zhou

What is the right way to reason about human activities? What directions forward are most promising? In this work, we analyze the current state of human activity understanding in videos. The goal of this paper is to examine datasets,…

计算机视觉与模式识别 · 计算机科学 2017-08-10 Gunnar A. Sigurdsson , Olga Russakovsky , Abhinav Gupta

Translating visual data into natural language is essential for machines to understand the world and interact with humans. In this work, a comprehensive study is conducted on video paragraph captioning, with the goal to generate…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Qinyu Li , Tengpeng Li , Hanli Wang , Chang Wen Chen

Tutorial videos of mobile apps have become a popular and compelling way for users to learn unfamiliar app features. To make the video accessible to the users, video creators always need to annotate the actions in the video, including what…

人机交互 · 计算机科学 2023-08-08 Sidong Feng , Chunyang Chen , Zhenchang Xing

Counting repetitive actions are widely seen in human activities such as physical exercise. Existing methods focus on performing repetitive action counting in short videos, which is tough for dealing with longer videos in more realistic…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Huazhang Hu , Sixun Dong , Yiqun Zhao , Dongze Lian , Zhengxin Li , Shenghua Gao

The temporal segmentation of events is an essential task and a precursor for the automatic recognition of human actions in the video. Several attempts have been made to capture frame-level salient aspects through attention but they lack the…

计算机视觉与模式识别 · 计算机科学 2020-05-08 Harshala Gammulle , Simon Denman , Sridha Sridharan , Clinton Fookes

Long-term activity forecasting is an especially challenging research problem because it requires understanding the temporal relationships between observed actions, as well as the variability and complexity of human activities. Despite…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Reuben Tan , Matthias De Lange , Michael Iuzzolino , Bryan A. Plummer , Kate Saenko , Karl Ridgeway , Lorenzo Torresani

Sequence transduction models have been widely explored in many natural language processing tasks. However, the target sequence usually consists of discrete tokens which represent word indices in a given vocabulary. We barely see the case…

计算机视觉与模式识别 · 计算机科学 2019-03-01 Xuan Liang , Yida Xu

In this paper we deal with the problem of predicting action progress in videos. We argue that this is an extremely important task since it can be valuable for a wide range of interaction applications. To this end we introduce a novel…

计算机视觉与模式识别 · 计算机科学 2020-03-11 Federico Becattini , Tiberio Uricchio , Lorenzo Seidenari , Lamberto Ballan , Alberto Del Bimbo

Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy…

机器人学 · 计算机科学 2025-12-22 Jonas Pai , Liam Achenbach , Victoriano Montesinos , Benedek Forrai , Oier Mees , Elvis Nava

In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos. Instructional videos contain complex activities and are a rich source of information for intelligent agents,…

计算机视觉与模式识别 · 计算机科学 2021-06-29 AJ Piergiovanni , Anelia Angelova , Michael S. Ryoo , Irfan Essa

Large-scale web-crawled datasets are fundamental for the success of pre-training vision-language models, such as CLIP. However, the inherent noise and potential irrelevance of web-crawled AltTexts pose challenges in achieving precise…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Zhengfeng Lai , Haotian Zhang , Bowen Zhang , Wentao Wu , Haoping Bai , Aleksei Timofeev , Xianzhi Du , Zhe Gan , Jiulong Shan , Chen-Nee Chuah , Yinfei Yang , Meng Cao

In this paper, we introduce a novel large-scale video dataset dubbed MM-SEAL for multi-person multi-grained spatio-temporal action localization among human daily life. We are the first to propose a new benchmark for multi-person…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Shimin Chen , Wei Li , Chen Chen , Jianyang Gu , Jiaming Chu , Xunqiang Tao , Yandong Guo

We explore the potential of large-scale generative video models for autonomous driving, introducing an open-source auto-regressive video model (VaViM) and its companion video-action model (VaVAM) to investigate how video pre-training…

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Kun Liu , Qi Liu , Xinchen Liu , Jie Li , Yongdong Zhang , Jiebo Luo , Xiaodong He , Wu Liu

Every moment counts in action recognition. A comprehensive understanding of human activity in video requires labeling every frame according to the actions occurring, placing multiple labels densely over a video sequence. To study this…

计算机视觉与模式识别 · 计算机科学 2017-06-12 Serena Yeung , Olga Russakovsky , Ning Jin , Mykhaylo Andriluka , Greg Mori , Li Fei-Fei

Understanding human motion beyond surface kinematics is crucial for motion analysis, rehabilitation, and injury risk assessment. However, progress in this domain is limited by the lack of large-scale datasets with biomechanical annotations,…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Yujun Huo , He Zhang , Chentao Song , Honglin Song , Zongyu Zuo , Tao Yu

This paper proposes Omni Dense Captioning, a novel task designed to generate continuous, fine-grained, and structured audio-visual narratives with explicit timestamps. To ensure dense semantic coverage, we introduce a six-dimensional…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Linli Yao , Yuancheng Wei , Yaojie Zhang , Lei Li , Xinlong Chen , Feifan Song , Ziyue Wang , Kun Ouyang , Yuanxin Liu , Lingpeng Kong , Qi Liu , Pengfei Wan , Kun Gai , Yuanxing Zhang , Xu Sun

Learning actions from human demonstration is an emerging trend for designing intelligent robotic systems, which can be referred as video to command. The performance of such approach highly relies on the quality of video captioning. However,…

计算机视觉与模式识别 · 计算机科学 2019-09-11 Shuo Yang , Wei Zhang , Weizhi Lu , Hesheng Wang , Yibin Li