中文
相关论文

相关论文: CaptainCook4D: A Dataset for Understanding Errors …

200 篇论文

Dietary intake estimation plays a crucial role in understanding the nutritional habits of individuals and populations, aiding in the prevention and management of diet-related health issues. Accurate estimation requires comprehensive…

计算机视觉与模式识别 · 计算机科学 2024-01-18 Chi-en Amy Tai , Saeejith Nair , Olivia Markham , Matthew Keller , Yifan Wu , Yuhao Chen , Alexander Wong

Many everyday tasks ranging from fixing appliances, cooking recipes to car maintenance require expert knowledge, especially when tasks are complex and multi-step. Despite growing interest in AI agents, there is a scarcity of dialogue-video…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Lavisha Aggarwal , Vikas Bahirwani , Lin Li , Andrea Colaco

We study multi-modal summarization for instructional videos, whose goal is to provide users an efficient way to learn skills in the form of text instructions and key video frames. We observe that existing benchmarks focus on generic…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Yuan Zang , Hao Tan , Seunghyun Yoon , Franck Dernoncourt , Jiuxiang Gu , Kushal Kafle , Chen Sun , Trung Bui

Information processing tasks involve complex cognitive mechanisms that are shaped by various factors, including individual goals, prior experience, and system environments. Understanding such behaviors requires a sophisticated and…

人机交互 · 计算机科学 2025-07-24 Kaixin Ji , Danula Hettiachchi , Falk Scholer , Flora D. Salim , Damiano Spina

Laparoscopic surgery is a complex surgical technique that requires extensive training. Recent advances in deep learning have shown promise in supporting this training by enabling automatic video-based assessment of surgical skills. However,…

We collected a new dataset that includes approximately eight hours of audiovisual recordings of a group of students and their self-evaluation scores for classroom engagement. The dataset and data analysis scripts are available on our…

人机交互 · 计算机科学 2023-04-19 Alpay Sabuncuoglu , T. Metin Sezgin

Most existing action quality assessment methods rely on the deep features of an entire video to predict the score, which is less reliable due to the non-transparent inference process and poor interpretability. We argue that understanding…

计算机视觉与模式识别 · 计算机科学 2022-04-08 Jinglin Xu , Yongming Rao , Xumin Yu , Guangyi Chen , Jie Zhou , Jiwen Lu

We study the problem of future step anticipation in procedural videos. Given a video of an ongoing procedural activity, we predict a plausible next procedure step described in rich natural language. While most previous work focus on the…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Mohamed Ashraf Abdelsalam , Samrudhdhi B. Rangrej , Isma Hadji , Nikita Dvornik , Konstantinos G. Derpanis , Afsaneh Fazly

Direct computer vision based-nutrient content estimation is a demanding task, due to deformation and occlusions of ingredients, as well as high intra-class and low inter-class variability between meal classes. In order to tackle these…

信息检索 · 计算机科学 2019-11-06 Matthias Fontanellaz , Stergios Christodoulidis , Stavroula Mougiakakou

Humans observe various actions being performed by other humans (physically or in videos/images) and can draw a wide range of inferences about it beyond what they can visually perceive. Such inferences include determining the aspects of the…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Shailaja Keyur Sampat , Yezhou Yang , Chitta Baral

Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI assistants. Effective guidance requires not only delivering…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Apratim Bhattacharyya , Bicheng Xu , Sanjay Haresh , Reza Pourreza , Litian Liu , Sunny Panchal , Pulkit Madan , Leonid Sigal , Roland Memisevic

This paper introduces a novel activity dataset which exhibits real-life and diverse scenarios of complex, temporally-extended human activities and actions. The dataset presents a set of videos of actors performing everyday activities in a…

计算机视觉与模式识别 · 计算机科学 2017-09-22 Jawad Tayyub , Majd Hawasly , David C. Hogg , Anthony G. Cohn

Guided troubleshooting is an inherent task in the domain of technical support services. When a customer experiences an issue with the functioning of a technical service or a product, an expert user helps guide the customer through a set of…

人工智能 · 计算机科学 2018-05-25 Abhirut Gupta , Abhay Khosla , Gautam Singh , Gargi Dasgupta

Human action recognition still exists many challenging problems such as different viewpoints, occlusion, lighting conditions, human body size and the speed of action execution, although it has been widely used in different areas. To tackle…

计算机视觉与模式识别 · 计算机科学 2024-06-10 Lei Wang

In this paper, we study the problem of procedure planning in instructional videos, which can be seen as a step towards enabling autonomous agents to plan for complex tasks in everyday settings such as cooking. Given the current visual…

计算机视觉与模式识别 · 计算机科学 2020-04-14 Chien-Yi Chang , De-An Huang , Danfei Xu , Ehsan Adeli , Li Fei-Fei , Juan Carlos Niebles

The fine-grained medical action analysis task has received considerable attention from pattern recognition communities recently, but it faces the problems of data and algorithm shortage. Cardiopulmonary Resuscitation (CPR) is an essential…

计算机视觉与模式识别 · 计算机科学 2023-09-22 Shunli Wang , Qing Yu , Shuaibing Wang , Dingkang Yang , Liuzhen Su , Xiao Zhao , Haopeng Kuang , Peixuan Zhang , Peng Zhai , Lihua Zhang

We present BASKET, a large-scale basketball video dataset for fine-grained skill estimation. BASKET contains 4,477 hours of video capturing 32,232 basketball players from all over the world. Compared to prior skill estimation datasets, our…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Yulu Pan , Ce Zhang , Gedas Bertasius

An adaptive guidance system that supports equipment operators requires a comprehensive model, which involves a variety of user behaviors that considers different skill and knowledge levels, as well as rapid-changing task situations. In the…

人机交互 · 计算机科学 2020-09-17 Chen Long-fei , Yuichi Nakamura , Kazuaki Kondo

Existing affective-computing, social-signal-processing, and meeting corpora capture important parts of human interaction, but they rarely support analysis of affect in co-located groups as a coupled individual, interpersonal, and…

Nowadays, it is common for people to take photographs of every beverage, snack, or meal they eat and then post these photographs on social media platforms. Leveraging these social trends, real-time food recognition and reliable…

计算机视觉与模式识别 · 计算机科学 2023-05-15 Aknur Karabay , Arman Bolatov , Huseyin Atakan Varol , Mei-Yen Chan