English
Related papers

Related papers: From Videos to Conversations: Egocentric Instructi…

200 papers

Many everyday tasks ranging from fixing appliances, cooking recipes to car maintenance require expert knowledge, especially when tasks are complex and multi-step. Despite growing interest in AI agents, there is a scarcity of dialogue-video…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Lavisha Aggarwal , Vikas Bahirwani , Lin Li , Andrea Colaco

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Keshigeyan Chandrasegaran , Agrim Gupta , Lea M. Hadzic , Taran Kota , Jimming He , Cristóbal Eyzaguirre , Zane Durante , Manling Li , Jiajun Wu , Li Fei-Fei

Comparing a user video to a reference how-to video is a key requirement for AR/VR technology delivering personalized assistance tailored to the user's progress. However, current approaches for language-based assistance can only answer…

Computer Vision and Pattern Recognition · Computer Science 2024-07-01 Tushar Nagarajan , Lorenzo Torresani

This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets, we propose an…

Computation and Language · Computer Science 2025-05-21 Yuxin Lin , Yinglin Zheng , Ming Zeng , Wangzheng Shi

While users tend to perceive instructional videos as an experience rather than a lesson with a set of instructions, instructional videos are more effective and appealing than textual user manuals and eliminate the ambiguity in text-based…

Human-Computer Interaction · Computer Science 2023-11-22 Songsong Liu , Shu Wang , Kun Sun

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an avatar chatbot system…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Se Jin Park , Chae Won Kim , Hyeongseop Rha , Minsu Kim , Joanna Hong , Jeong Hun Yeo , Yong Man Ro

Recent advances in conversational AI have been substantial, but developing real-time systems for perceptual task guidance remains challenging. These systems must provide interactive, proactive assistance based on streaming visual inputs,…

Artificial Intelligence · Computer Science 2025-06-09 Yichi Zhang , Xin Luna Dong , Zhaojiang Lin , Andrea Madotto , Anuj Kumar , Babak Damavandi , Joyce Chai , Seungwhan Moon

The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective…

Artificial Intelligence · Computer Science 2026-03-03 Hengjian Gao , Kaiwei Zhang , Shibo Wang , Mingjie Chen , Qihang Cao , Xianfeng Wang , Yucheng Zhu , Xiongkuo Min , Wei Sun , Dandan Zhu , Guangtao Zhai

A well-designed interactive human-like dialogue system is expected to take actions (e.g. smiling) and respond in a pattern similar to humans. However, due to the limitation of single-modality (only speech) or small volume of currently…

Human-Computer Interaction · Computer Science 2022-12-13 Zhiling Luo , Qiankun Shi , Sha Zhao , Wei Zhou , Haiqing Chen , Yuankai Ma , Haitao Leng

It is still a pipe dream that personal AI assistants on the phone and AR glasses can assist our daily life in addressing our questions like ``how to adjust the date for this watch?'' and ``how to set its heating duration? (while pointing at…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Stan Weixian Lei , Difei Gao , Yuxuan Wang , Dongxing Mao , Zihan Liang , Lingmin Ran , Mike Zheng Shou

In this study, conversations between humans and avatars are linguistically, organizationally, and structurally analyzed, focusing on what is necessary for creating face-to-face multimodal interfaces for machines. We videorecorded…

Human-Computer Interaction · Computer Science 2022-11-28 João Ranhel , Cacilda Vilela de Lima

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Liang Xu , Chengqun Yang , Zili Lin , Fei Xu , Yifan Liu , Congsheng Xu , Yiyi Zhang , Jie Qin , Xingdong Sheng , Yunhui Liu , Xin Jin , Yichao Yan , Wenjun Zeng , Xiaokang Yang

Communicating in noisy, multi-talker environments is challenging, especially for people with hearing impairments. Egocentric video data can potentially be used to identify a user's conversation partners, which could be used to inform…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Tobias Dorszewski , Søren A. Fuglsang , Jens Hjortkjær

Learning text-video embeddings usually requires a dataset of video clips with manually provided captions. However, such datasets are expensive and time consuming to create and therefore difficult to obtain on a large scale. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2019-08-01 Antoine Miech , Dimitri Zhukov , Jean-Baptiste Alayrac , Makarand Tapaswi , Ivan Laptev , Josef Sivic

We introduce a new task -- language-driven video inpainting, which uses natural language instructions to guide the inpainting process. This approach overcomes the limitations of traditional video inpainting methods that depend on manually…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Jianzong Wu , Xiangtai Li , Chenyang Si , Shangchen Zhou , Jingkang Yang , Jiangning Zhang , Yining Li , Kai Chen , Yunhai Tong , Ziwei Liu , Chen Change Loy

People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our observations of visual rehabilitation therapists (VRTs)…

Human-Computer Interaction · Computer Science 2025-07-28 Mina Huh , Zihui Xue , Ujjaini Das , Kumar Ashutosh , Kristen Grauman , Amy Pavel

Existing studies on talking video generation have predominantly focused on single-person monologues or isolated facial animations, limiting their applicability to realistic multi-human interactions. To bridge this gap, we introduce MIT, a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Zeyu Zhu , Weijia Wu , Mike Zheng Shou

A long-standing goal of intelligent assistants such as AR glasses/robots has been to assist users in affordance-centric real-world scenarios, such as "how can I run the microwave for 1 minute?". However, there is still no clear task…

Computer Vision and Pattern Recognition · Computer Science 2022-07-21 Benita Wong , Joya Chen , You Wu , Stan Weixian Lei , Dongxing Mao , Difei Gao , Mike Zheng Shou

Intelligent assistance involves not only understanding but also action. Existing ego-centric video datasets contain rich annotations of the videos, but not of actions that an intelligent assistant could perform in the moment. To address…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Steven Abreu , Tiffany D. Do , Karan Ahuja , Eric J. Gonzalez , Lee Payne , Daniel McDuff , Mar Gonzalez-Franco

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual…

Computer Vision and Pattern Recognition · Computer Science 2021-08-13 Antoine Yang , Antoine Miech , Josef Sivic , Ivan Laptev , Cordelia Schmid
‹ Prev 1 2 3 10 Next ›