English

From Videos to Conversations: Egocentric Instructions for Task Assistance

Computer Vision and Pattern Recognition 2026-02-03 v1

Abstract

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (AR) assistance, progress remains limited by the scarcity of large-scale multimodal conversational datasets grounded in real-world task execution, in part due to the cost and logistical complexity of human-assisted data collection. In this paper, we present a framework to automatically transform single person instructional videos into two-person multimodal task-guidance conversations. Our fully automatic pipeline, based on large language models, provides a scalable and cost efficient alternative to traditional data collection approaches. Using this framework, we introduce HowToDIV, a multimodal dataset comprising 507 conversations, 6,636 question answer pairs, and 24 hours of video spanning multiple domains. Each session consists of a multi-turn expert-novice interaction. Finally, we report baseline results using Gemma 3 and Qwen 2.5 on HowToDIV, providing an initial benchmark for multimodal procedural task assistance.

Keywords

Cite

@article{arxiv.2602.01038,
  title  = {From Videos to Conversations: Egocentric Instructions for Task Assistance},
  author = {Lavisha Aggarwal and Vikas Bahirwani and Andrea Colaco},
  journal= {arXiv preprint arXiv:2602.01038},
  year   = {2026}
}
R2 v1 2026-07-01T09:29:54.647Z