中文
相关论文

相关论文: CaptainCook4D: A Dataset for Understanding Errors …

200 篇论文

This paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS. The effort culminates in EPIC-KITCHENS-100, a collection of 100 hours, 20M frames, 90K actions in 700 variable-length videos, capturing…

Mistake analysis in procedural activities is a critical area of research with applications spanning industrial automation, physical rehabilitation, education and human-robot collaboration. This paper reviews vision-based methods for…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Konstantinos Bacharidis , Antonis A. Argyros

Although action recognition for procedural tasks has received notable attention, it has a fundamental flaw in that no measure of success for actions is provided. This limits the applicability of such systems especially within the industrial…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Tim J. Schoonbeek , Tim Houben , Hans Onvlee , Peter H. N. de With , Fons van der Sommen

Despite the recent progress on 6D object pose estimation methods for robotic grasping, a substantial performance gap persists between the capabilities of these methods on existing datasets and their efficacy in real-world grasping and…

机器人学 · 计算机科学 2024-12-18 Abdelrahman Younes , Tamim Asfour

Procedure learning involves identifying the key-steps and determining their logical order to perform a task. Existing approaches commonly use third-person videos for learning the procedure, making the manipulated object small in appearance…

计算机视觉与模式识别 · 计算机科学 2022-07-25 Siddhant Bansal , Chetan Arora , C. V. Jawahar

Understanding the complexity of human activities solely through an individual's data can be challenging. However, in many situations, surrounding individuals are likely performing similar activities, while existing human activity…

计算机视觉与模式识别 · 计算机科学 2023-06-12 Haoxiang Yu , Jingyi An , Evan King , Edison Thomaz , Christine Julien

Computer vision has a great potential to help our daily lives by searching for lost keys, watering flowers or reminding us to take a pill. To succeed with such tasks, computer vision methods need to be trained from real and diverse examples…

计算机视觉与模式识别 · 计算机科学 2016-07-28 Gunnar A. Sigurdsson , Gül Varol , Xiaolong Wang , Ali Farhadi , Ivan Laptev , Abhinav Gupta

Action recognition is so far mainly focusing on the problem of classification of hand selected preclipped actions and reaching impressive results in this field. But with the performance even ceiling on current datasets, it also appears that…

计算机视觉与模式识别 · 计算机科学 2019-06-05 Hilde Kuehne , Ahsan Iqbal , Alexander Richard , Juergen Gall

Procedures are inherently hierarchical. To "make videos", one may need to "purchase a camera", which in turn may require one to "set a budget". While such hierarchical knowledge is critical for reasoning about complex procedures, most…

计算与语言 · 计算机科学 2022-03-18 Shuyan Zhou , Li Zhang , Yue Yang , Qing Lyu , Pengcheng Yin , Chris Callison-Burch , Graham Neubig

Recent approaches in depth-based human activity analysis achieved outstanding performance and proved the effectiveness of 3D representation for classification of action classes. Currently available depth-based and RGB+D-based action…

计算机视觉与模式识别 · 计算机科学 2016-04-12 Amir Shahroudy , Jun Liu , Tian-Tsong Ng , Gang Wang

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

计算机视觉与模式识别 · 计算机科学 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

Understanding animals' behaviors is significant for a wide range of applications. However, existing animal behavior datasets have limitations in multiple aspects, including limited numbers of animal classes, data samples and provided tasks,…

计算机视觉与模式识别 · 计算机科学 2022-06-06 Xun Long Ng , Kian Eng Ong , Qichen Zheng , Yun Ni , Si Yong Yeo , Jun Liu

There are substantial instructional videos on the Internet, which enables us to acquire knowledge for completing various tasks. However, most existing datasets for instructional video analysis have the limitations in diversity and…

计算机视觉与模式识别 · 计算机科学 2019-03-08 Yansong Tang , Dajun Ding , Yongming Rao , Yu Zheng , Danyang Zhang , Lili Zhao , Jiwen Lu , Jie Zhou

Understanding food recipe requires anticipating the implicit causal effects of cooking actions, such that the recipe can be converted into a graph describing the temporal workflow of the recipe. This is a non-trivial task that involves…

计算与语言 · 计算机科学 2020-08-24 Liangming Pan , Jingjing Chen , Jianlong Wu , Shaoteng Liu , Chong-Wah Ngo , Min-Yen Kan , Yu-Gang Jiang , Tat-Seng Chua

Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resources, making it…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Koki Maeda , Tosho Hirasawa , Atsushi Hashimoto , Jun Harashima , Leszek Rybicki , Yusuke Fukasawa , Yoshitaka Ushiku

Understanding procedural texts, such as cooking recipes, is essential for enabling machines to follow instructions and reason about tasks, a key aspect of intelligent reasoning. In cooking, these instructions can be interpreted as a series…

计算与语言 · 计算机科学 2024-10-11 Aissatou Diallo , Antonis Bikakis , Luke Dickens , Anthony Hunter , Rob Miller

In the development of science, accurate and reproducible documentation of the experimental process is crucial. Automatic recognition of the actions in experiments from videos would help experimenters by complementing the recording of…

计算机视觉与模式识别 · 计算机科学 2024-02-02 Takuma Yagi , Misaki Ohashi , Yifei Huang , Ryosuke Furuta , Shungo Adachi , Toutai Mitsuyama , Yoichi Sato

Understanding how humans cooperatively rearrange household objects is critical for VR/AR and human-robot interaction. However, in-depth studies on modeling these behaviors are under-researched due to the lack of relevant datasets. We fill…

计算机视觉与模式识别 · 计算机科学 2024-12-25 Yun Liu , Chengwen Zhang , Ruofan Xing , Bingda Tang , Bowen Yang , Li Yi

We present a validation dataset of newly-collected kitchen-based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional…

People interacting with voice assistants are often frustrated by voice assistants' frequent errors and inability to respond to backchannel cues. We introduce an open-source video dataset of 21 participants' interactions with a voice…

人机交互 · 计算机科学 2021-04-16 Andrea Cuadra , Hansol Lee , Jason Cho , Wendy Ju