English
Related papers

Related papers: Assembly101: A Large-Scale Multi-View Video Datase…

200 papers

We present AssemblyHands, a large-scale benchmark dataset with accurate 3D hand pose annotations, to facilitate the study of egocentric activities with challenging hand-object interactions. The dataset includes synchronized egocentric and…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Takehiko Ohkawa , Kun He , Fadime Sener , Tomas Hodan , Luan Tran , Cem Keskin

With the emergence of collaborative robots (cobots), human-robot collaboration in industrial manufacturing is coming into focus. For a cobot to act autonomously and as an assistant, it must understand human actions during assembly. To…

Robotics · Computer Science 2023-04-18 Dustin Aganian , Benedict Stephan , Markus Eisenbach , Corinna Stretz , Horst-Michael Gross

Understanding comprehensive assembly knowledge from videos is critical for futuristic ultra-intelligent industry. To enable technological breakthrough, we present HA-ViD - the first human assembly video dataset that features representative…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Hao Zheng , Regina Lee , Yuqian Lu

This work presents the Industrial Hand Action Dataset V1, an industrial assembly dataset consisting of 12 classes with 459,180 images in the basic version and 2,295,900 images after spatial augmentation. Compared to other freely available…

Computer Vision and Pattern Recognition · Computer Science 2023-03-08 Fabian Sturm , Elke Hergenroether , Julian Reinhardt , Petar Smilevski Vojnovikj , Melanie Siegel

Visual segmentation has seen tremendous advancement recently with ready solutions for a wide variety of scene types, including human hands and other body parts. However, focus on segmentation of human hands while performing complex tasks,…

Computer Vision and Pattern Recognition · Computer Science 2019-02-21 Roy Shilkrot , Zhi Chai , Minh Hoai

Robotic manipulation remains a core challenge in robotics, particularly for contact-rich tasks such as industrial assembly and disassembly. Existing datasets have significantly advanced learning in manipulation but are primarily focused on…

Inferring physical actions from visual observations is a fundamental capability for advancing machine intelligence in the physical world. Achieving this requires large-scale, open-vocabulary video action datasets that span broad domains. We…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Delong Chen , Tejaswi Kasarla , Yejin Bang , Mustafa Shukor , Willy Chung , Jade Yu , Allen Bolourchi , Theo Moutakanni , Pascale Fung

Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly step. Existing datasets focus on simplified scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Danrui Li , Jiahao Zhang , Bernhard Egger , Moitreya Chatterjee , Suhas Lohit , Tim K. Marks , Anoop Cherian

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions…

Computer Vision and Pattern Recognition · Computer Science 2019-02-19 Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan , Lisa Brown , Quanfu Fan , Dan Gutfruend , Carl Vondrick , Aude Oliva

We introduce UCF101 which is currently the largest dataset of human actions. It consists of 101 action classes, over 13k clips and 27 hours of video data. The database consists of realistic user uploaded videos containing camera motion and…

Computer Vision and Pattern Recognition · Computer Science 2012-12-04 Khurram Soomro , Amir Roshan Zamir , Mubarak Shah

Bimanual human activities inherently involve coordinated movements of both hands and body. However, the impact of this coordination in activity understanding has not been systematically evaluated due to the lack of suitable datasets. Such…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Tatsuro Banno , Takehiko Ohkawa , Ruicong Liu , Ryosuke Furuta , Yoichi Sato

In recent years, we have seen an emergence of data-driven approaches in robotics. However, most existing efforts and datasets are either in simulation or focus on a single task in isolation such as grasping, pushing or poking. In order to…

Robotics · Computer Science 2018-10-17 Pratyusha Sharma , Lekha Mohan , Lerrel Pinto , Abhinav Gupta

A new large-scale video dataset for human action recognition, called STAIR Actions is introduced. STAIR Actions contains 100 categories of action labels representing fine-grained everyday home actions so that it can be applied to research…

Computer Vision and Pattern Recognition · Computer Science 2018-04-17 Yuya Yoshikawa , Jiaqing Lin , Akikazu Takeuchi

In this paper we address the task of recognizing assembly actions as a structure (e.g. a piece of furniture or a toy block tower) is built up from a set of primitive objects. Recognizing the full range of assembly actions requires…

Computer Vision and Pattern Recognition · Computer Science 2020-12-03 Jonathan D. Jones , Cathryn Cortesa , Amy Shelton , Barbara Landau , Sanjeev Khudanpur , Gregory D. Hager

Video action recognition is one of the representative tasks for video understanding. Over the last decade, we have witnessed great advancements in video action recognition thanks to the emergence of deep learning. But we also encountered…

Computer Vision and Pattern Recognition · Computer Science 2020-12-14 Yi Zhu , Xinyu Li , Chunhui Liu , Mohammadreza Zolfaghari , Yuanjun Xiong , Chongruo Wu , Zhi Zhang , Joseph Tighe , R. Manmatha , Mu Li

Human actions in videos are 3D signals. However, there are a few methods available for multiple human action recognition. For long videos, it's difficult to search within a video for a specific action and/or person. For that, this paper…

Computer Vision and Pattern Recognition · Computer Science 2021-03-16 Noor Almaadeed , Omar Elharrouss , Somaya Al-Maadeed , Ahmed Bouridane , Azeddine Beghdadi

Fine-grained understanding of human actions and poses in videos is essential for human-centric AI applications. In this work, we introduce ActionArt, a fine-grained video-caption dataset designed to advance research in human-centric…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Yi-Xing Peng , Qize Yang , Yu-Ming Tang , Shenghao Fu , Kun-Yu Lin , Xihan Wei , Wei-Shi Zheng

In online action detection, the goal is to detect the start of an action in a video stream as soon as it happens. For instance, if a child is chasing a ball, an autonomous car should recognize what is going on and respond immediately. This…

Computer Vision and Pattern Recognition · Computer Science 2016-08-31 Roeland De Geest , Efstratios Gavves , Amir Ghodrati , Zhenyang Li , Cees Snoek , Tinne Tuytelaars

Advancements in deep neural networks have contributed to near perfect results for many computer vision problems such as object recognition, face recognition and pose estimation. However, human action recognition is still far from…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Asanka G. Perera , Yee Wei Law , Titilayo T. Ogunwa , Javaan Chahl

We present WayveScenes101, a dataset designed to help the community advance the state of the art in novel view synthesis that focuses on challenging driving scenes containing many dynamic and deformable elements with changing geometry and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Jannik Zürn , Paul Gladkov , Sofía Dudas , Fergal Cotter , Sofi Toteva , Jamie Shotton , Vasiliki Simaiaki , Nikhil Mohan
‹ Prev 1 2 3 10 Next ›