English
Related papers

Related papers: HA-ViD: A Human Assembly Video Dataset for Compreh…

200 papers

With the emergence of collaborative robots (cobots), human-robot collaboration in industrial manufacturing is coming into focus. For a cobot to act autonomously and as an assistant, it must understand human actions during assembly. To…

Robotics · Computer Science 2023-04-18 Dustin Aganian , Benedict Stephan , Markus Eisenbach , Corinna Stretz , Horst-Michael Gross

We present HumanEdit, a high-quality, human-rewarded dataset specifically designed for instruction-guided image editing, enabling precise and diverse image manipulations through open-form language instructions. Previous large-scale editing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Jinbin Bai , Wei Chow , Ling Yang , Xiangtai Li , Juncheng Li , Hanwang Zhang , Shuicheng Yan

Detecting and interpreting operator actions, engagement, and object interactions in dynamic industrial workflows remains a significant challenge in human-robot collaboration research, especially within complex, real-world environments.…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Naval Kishore Mehta , Arvind , Himanshu Kumar , Abeer Banerjee , Sumeet Saurav , Sanjay Singh

This paper introduces a vision-based framework for capturing and understanding human behavior in industrial assembly lines, focusing on car door manufacturing. The framework leverages advanced computer vision techniques to estimate workers'…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Konstantinos Papoutsakis , Nikolaos Bakalos , Konstantinos Fragkoulis , Athena Zacharia , Georgia Kapetadimitri , Maria Pateraki

This paper introduces a new video-and-language dataset with human actions for multimodal logical inference, which focuses on intentional and aspectual expressions that describe dynamic human actions. The dataset consists of 200 videos,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Riko Suzuki , Hitomi Yanaka , Koji Mineshima , Daisuke Bekki

Mimicry is a fundamental learning mechanism in humans, enabling individuals to learn new tasks by observing and imitating experts. However, applying this ability to robots presents significant challenges due to the inherent differences…

Robotics · Computer Science 2025-09-23 Hanjung Kim , Jaehyun Kang , Hyolim Kang , Meedeum Cho , Seon Joo Kim , Youngwoon Lee

We present AVID, the first large-scale benchmark for audio-visual inconsistency understanding in videos. While omni-modal large language models excel at temporally aligned tasks such as captioning and question answering, they struggle to…

Multimedia · Computer Science 2026-04-16 Zixuan Chen , Depeng Wang , Hao Lin , Li Luo , Ke Xu , Ya Guo , Huijia Zhu , Tanfeng Sun , Xinghao Jiang

Current video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning disobeys the natural perception that humans do in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-04 Khoa Vo , Thinh Phan , Kashu Yamazaki , Minh Tran , Ngan Le

Humans detect real-world object anomalies by perceiving, interacting, and reasoning based on object-conditioned physical knowledge. The long-term goal of Industrial Anomaly Detection (IAD) is to enable machines to autonomously replicate…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Wenqiao Li , Yao Gu , Xintao Chen , Xiaohao Xu , Ming Hu , Xiaonan Huang , Yingna Wu

We present a domain- and user-preference-agnostic approach to detect highlightable excerpts from human-centric videos. Our method works on the graph-based representation of multiple observable human-centric modalities in the videos, such as…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Uttaran Bhattacharya , Gang Wu , Stefano Petrangeli , Viswanathan Swaminathan , Dinesh Manocha

Tutorial videos are a valuable resource for people looking to learn new tasks. People often learn these skills by viewing multiple tutorial videos to get an overall understanding of a task by looking at different approaches to achieve the…

Human-Computer Interaction · Computer Science 2025-03-28 Saelyne Yang , Anh Truong , Juho Kim , Dingzeyu Li

Human parsing aims to partition humans in image or video into multiple pixel-level semantic parts. In the last decade, it has gained significantly increased interest in the computer vision community and has been utilized in a broad range of…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Lu Yang , Wenhe Jia , Shan Li , Qing Song

Generic motion understanding from video involves not only tracking objects, but also perceiving how their surfaces deform and move. This information is useful to make inferences about 3D shape, physical properties and object interactions.…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Carl Doersch , Ankush Gupta , Larisa Markeeva , Adrià Recasens , Lucas Smaira , Yusuf Aytar , João Carreira , Andrew Zisserman , Yi Yang

The advancement of embodied AI has unlocked significant potential for intelligent humanoid robots. However, progress in both Vision-Language-Action (VLA) models and world models is severely hampered by the scarcity of large-scale, diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Pei Yang , Hai Ci , Yiren Song , Mike Zheng Shou

Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes. A complex event relates to dense crowds, anomalous individuals, or…

Computer Vision and Pattern Recognition · Computer Science 2023-07-14 Weiyao Lin , Huabin Liu , Shizhan Liu , Yuxi Li , Rui Qian , Tao Wang , Ning Xu , Hongkai Xiong , Guo-Jun Qi , Nicu Sebe

Most existing video tasks related to "human" focus on the segmentation of salient humans, ignoring the unspecified others in the video. Few studies have focused on segmenting and tracking all humans in a complex video, including pedestrians…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Ran Yu , Chenyu Tian , Weihao Xia , Xinyuan Zhao , Haoqian Wang , Yujiu Yang

In this paper we introduce a new dataset for 360-degree video summarization: the transformation of 360-degree video content to concise 2D-video summaries that can be consumed via traditional devices, such as TV sets and smartphones. The…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Ioannis Kontostathis , Evlampios Apostolidis , Vasileios Mezaris

Visual segmentation has seen tremendous advancement recently with ready solutions for a wide variety of scene types, including human hands and other body parts. However, focus on segmentation of human hands while performing complex tasks,…

Computer Vision and Pattern Recognition · Computer Science 2019-02-21 Roy Shilkrot , Zhi Chai , Minh Hoai

Understanding how humans interact with each other is key to building realistic multi-human virtual reality systems. This area remains relatively unexplored due to the lack of large-scale datasets. Recent datasets focusing on this issue…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Rawal Khirodkar , Jyun-Ting Song , Jinkun Cao , Zhengyi Luo , Kris Kitani

In human vision objects and their parts can be visually recognized from purely spatial or purely temporal information but the mechanisms integrating space and time are poorly understood. Here we show that human visual recognition of objects…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Guy Ben-Yosef , Gabriel Kreiman , Shimon Ullman