中文
相关论文

相关论文: CaptainCook4D: A Dataset for Understanding Errors …

200 篇论文

With the continuously thriving popularity around the world, fitness activity analytic has become an emerging research topic in computer vision. While a variety of new tasks and algorithms have been proposed recently, there are growing…

计算机视觉与模式识别 · 计算机科学 2023-04-20 Yansong Tang , Jinpeng Liu , Aoyang Liu , Bin Yang , Wenxun Dai , Yongming Rao , Jiwen Lu , Jie Zhou , Xiu Li

Imitation learning from large multi-task demonstration datasets has emerged as a promising path for building generally-capable robots. As a result, 1000s of hours have been spent on building such large-scale datasets around the globe.…

Secondary analysis or the reuse of existing survey data is a common practice among social scientists. Searching for relevant datasets in Digital Libraries is a somehow unfamiliar behaviour for this community. Dataset retrieval, especially…

数字图书馆 · 计算机科学 2020-10-13 Zeljko Carevic , Dwaipayan Roy , Philipp Mayr

Human Action Recognition (HAR) is a very crucial task in computer vision. It helps to carry out a series of downstream tasks, like understanding human behaviors. Due to the complexity of human behaviors, many highly valuable behaviors are…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Hongwu Li , Zhenliang Zhang , Wei Wang

To train deep learning models for vision-based action recognition of elders' daily activities, we need large-scale activity datasets acquired under various daily living environments and conditions. However, most public datasets used in…

计算机视觉与模式识别 · 计算机科学 2020-11-22 Hochul Hwang , Cheongjae Jang , Geonwoo Park , Junghyun Cho , Ig-Jae Kim

Engagement in virtual learning is essential for participant satisfaction, performance, and adherence, particularly in online education and virtual rehabilitation, where interactive communication plays a key role. Yet, accurately measuring…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Ali Abedi , Sadaf Safa , Tracey J. F. Colella , Shehroz S. Khan

We introduce the task of early mistake detection in video, where the goal is to determine whether a keystep in a procedural activity is performed correctly while observing as little of the streaming video as possible. To tackle this…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Sagnik Majumder , Anish Nethi , Ziad Al-Halah , Kristen Grauman

Imitation learning for manipulation has a well-known data scarcity problem. Unlike natural language and 2D computer vision, there is no Internet-scale corpus of data for dexterous manipulation. One appealing option is egocentric human…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Ryan Hoque , Peide Huang , David J. Yoon , Mouli Sivapurapu , Jian Zhang

Existing activity tracker datasets for human activity recognition are typically obtained by having participants perform predefined activities in an enclosed environment under supervision. This results in small datasets with a limited number…

人机交互 · 计算机科学 2024-03-01 Shing Chan , Hang Yuan , Catherine Tong , Aidan Acquah , Abram Schonfeldt , Jonathan Gershuny , Aiden Doherty

In this paper, we are interested in modeling a how-to instructional procedure, such as a cooking recipe, with a meaningful and rich high-level representation. Specifically, we propose to represent cooking recipes and food images as cooking…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Dim P. Papadopoulos , Enrique Mora , Nadiia Chepurko , Kuan Wei Huang , Ferda Ofli , Antonio Torralba

To help accelerate progress in multi-target, multi-camera tracking systems, we present (i) a new pair of precision-recall measures of performance that treats errors of all types uniformly and emphasizes correct identification over sources…

计算机视觉与模式识别 · 计算机科学 2016-09-20 Ergys Ristani , Francesco Solera , Roger S. Zou , Rita Cucchiara , Carlo Tomasi

We present EgoFun3D, a coordinated task formulation, dataset, and benchmark for modeling interactive 3D objects from egocentric videos. Interactive objects are of high interest for embodied AI but scarce, making modeling from readily…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Weikun Peng , Denys Iliash , Manolis Savva

In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this model remains static for new users and environments, we…

计算机视觉与模式识别 · 计算机科学 2023-07-13 Matthias De Lange , Hamid Eghbalzadeh , Reuben Tan , Michael Iuzzolino , Franziska Meier , Karl Ridgeway

We propose a dataset to study the influence of object-specific characteristics on human pick-and-place movements and compare the quality of the motion kinematics extracted by various sensors. This dataset is also suitable for promoting a…

Computer-assisted minimally invasive surgery has great potential in benefiting modern operating theatres. The video data streamed from the endoscope provides rich information to support context-awareness for next-generation intelligent…

计算机视觉与模式识别 · 计算机科学 2022-08-04 Ziyi Wang , Bo Lu , Yonghao Long , Fangxun Zhong , Tak-Hong Cheung , Qi Dou , Yunhui Liu

Commonsense procedural knowledge is important for AI agents and robots that operate in a human environment. While previous attempts at constructing procedural knowledge are mostly rule- and template-based, recent advances in deep learning…

计算与语言 · 计算机科学 2019-09-17 Yilun Zhou , Julie A. Shah , Steven Schockaert

In this paper, we are interested in modeling complex activities that occur in a typical household. We propose to use programs, i.e., sequences of atomic actions and interactions, as a high level representation of complex tasks. Programs are…

计算机视觉与模式识别 · 计算机科学 2018-06-20 Xavier Puig , Kevin Ra , Marko Boben , Jiaman Li , Tingwu Wang , Sanja Fidler , Antonio Torralba

Procedural videos, exemplified by recipe demonstrations, are instrumental in conveying step-by-step instructions. However, understanding such videos is challenging as it involves the precise localization of steps and the generation of…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Anil Batra , Davide Moltisanti , Laura Sevilla-Lara , Marcus Rohrbach , Frank Keller

In recent years, the landscape of computer-assisted interventions and post-operative surgical video analysis has been dramatically reshaped by deep-learning techniques, resulting in significant advancements in surgeons' skills, operation…

Procedural knowledge describes how to accomplish tasks and mitigate problems. Such knowledge is commonly held by domain experts, e.g. operators in manufacturing who adjust parameters to achieve quality targets. To the best of our knowledge,…

人工智能 · 计算机科学 2023-08-17 Richard Nordsieck , André Schweizer , Michael Heider , Jörg Hähner