English
Related papers

Related papers: Assembly101: A Large-Scale Multi-View Video Datase…

200 papers

We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 735 sequences (7,875 seconds at 30 fps), HUMOTO captures…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Jiaxin Lu , Chun-Hao Paul Huang , Uttaran Bhattacharya , Qixing Huang , Yi Zhou

Part-level Action Parsing aims at part state parsing for boosting action recognition in videos. Despite of dramatic progresses in the area of video classification research, a severe problem faced by the community is that the detailed…

Computer Vision and Pattern Recognition · Computer Science 2021-11-08 Xuanhan Wang , Xiaojia Chen , Lianli Gao , Lechao Chen , Jingkuan Song

Despite the fact that notable improvements have been made recently in the field of feature extraction and classification, human action recognition is still challenging, especially in images, in which, unlike videos, there is no motion.…

Computer Vision and Pattern Recognition · Computer Science 2020-03-24 Sina Mohammadi , Sina Ghofrani Majelan , Shahriar B. Shokouhi

This paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage both consensus and disagreement among…

Computer Vision and Pattern Recognition · Computer Science 2019-09-05 Hang Zhao , Antonio Torralba , Lorenzo Torresani , Zhicheng Yan

Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Yufan Deng , Daquan Zhou

The advancement of computer vision and machine learning has made datasets a crucial element for further research and applications. However, the creation and development of robots with advanced recognition capabilities are hindered by the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-24 Zhengcheng Shen , Yi Gao , Linh Kästner , Jens Lambrecht

This paper proposes a robot assembly planning method by automatically reading the graphical instruction manuals design for humans. Essentially, the method generates an Assembly Task Sequence Graph (ATSG) by recognizing a graphical…

This paper develops a planner to find an optimal assembly sequence to assemble several objects. The input to the planner is the mesh models of the objects, the relative poses between the objects in the assembly, and the final pose of the…

Robotics · Computer Science 2016-09-13 Weiwei Wan , Kensuke Harada , Kazuyuki Nagata

Current perception models in autonomous driving have become notorious for greatly relying on a mass of annotated data to cover unseen cases and address the long-tail problem. On the other hand, learning from unlabeled large-scale collected…

Computer Vision and Pattern Recognition · Computer Science 2021-10-26 Jiageng Mao , Minzhe Niu , Chenhan Jiang , Hanxue Liang , Jingheng Chen , Xiaodan Liang , Yamin Li , Chaoqiang Ye , Wei Zhang , Zhenguo Li , Jie Yu , Hang Xu , Chunjing Xu

The scarcity of high quality actions video data is a bottleneck in the research and application of action recognition. Although significant effort has been made in this area, there still exist gaps in the range of available data types a…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Shuo Wang , Amiya Ranjan , Lawrence Jiang

Determining when people are struggling allows for a finer-grained understanding of actions that complements conventional action classification and error detection. Struggle detection, as defined in this paper, is a distinct and important…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Shijia Feng , Michael Wray , Brian Sullivan , Youngkyoon Jang , Casimir Ludwig , Iain Gilchrist , Walterio Mayol-Cuevas

The temporal action segmentation task segments videos temporally and predicts action labels for all frames. Fully supervising such a segmentation model requires dense frame-wise action annotations, which are expensive and tedious to…

Computer Vision and Pattern Recognition · Computer Science 2022-12-20 Guodong Ding , Angela Yao

Recently, much progress has been made for self-supervised action recognition. Most existing approaches emphasize the contrastive relations among videos, including appearance and motion consistency. However, two main issues remain for…

Computer Vision and Pattern Recognition · Computer Science 2022-04-28 Guanhong Wang , Keyu Lu , Yang Zhou , Zhanhao He , Gaoang Wang

In this paper, we introduce a new hierarchical model for human action recognition using body joint locations. Our model can categorize complex actions in videos, and perform spatio-temporal annotations of the atomic actions that compose the…

Computer Vision and Pattern Recognition · Computer Science 2016-06-17 Ivan Lillo , Juan Carlos Niebles , Alvaro Soto

Multi-agent behavior modeling aims to understand the interactions that occur between agents. We present a multi-agent dataset from behavioral neuroscience, the Caltech Mouse Social Interactions (CalMS21) Dataset. Our dataset consists of…

Human actions often involve complex interactions across several inter-related objects in the scene. However, existing approaches to fine-grained video understanding or visual relationship detection often rely on single object representation…

Computer Vision and Pattern Recognition · Computer Science 2018-03-22 Chih-Yao Ma , Asim Kadav , Iain Melvin , Zsolt Kira , Ghassan AlRegib , Hans Peter Graf

Automated monitoring of construction operations, especially operations of equipment and machines, is an essential step toward cost-estimating, and planning of construction projects. In recent years, a number of methods were suggested for…

Signal Processing · Electrical Eng. & Systems 2019-06-06 Behnam Sherafat , Abbas Rashidi , Yong-Cheol Lee , Changbum R. Ahn

Modelling interactions between humans and objects in natural environments is central to many applications including gaming, virtual and mixed reality, as well as human behavior analysis and human-robot collaboration. This challenging…

Computer Vision and Pattern Recognition · Computer Science 2022-04-15 Bharat Lal Bhatnagar , Xianghui Xie , Ilya A. Petrov , Cristian Sminchisescu , Christian Theobalt , Gerard Pons-Moll

Action recognition in videos is a challenging task due to the complexity of the spatio-temporal patterns to model and the difficulty to acquire and learn on large quantities of video data. Deep learning, although a breakthrough for image…

Computer Vision and Pattern Recognition · Computer Science 2016-08-26 César Roberto de Souza , Adrien Gaidon , Eleonora Vig , Antonio Manuel López

The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or potentially even disabled users. We call this task…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Alejandro Pardo , Jui-Hsien Wang , Bernard Ghanem , Josef Sivic , Bryan Russell , Fabian Caba Heilbron