中文
相关论文

相关论文: Vid2Coach: Transforming How-To Videos into Task As…

200 篇论文

This work addresses challenges in developing conversational assistants that support rich multimodal video interactions to accomplish real-world tasks interactively. We introduce the task of automatically linking instructional videos to task…

信息检索 · 计算机科学 2022-08-24 Sophie Fischer , Carlos Gemmell , Iain Mackie , Jeffrey Dalton

Videos offer rich audiovisual information that can support people in performing activities of daily living (ADLs), but they remain largely inaccessible to blind or low-vision (BLV) individuals. In cooking, BLV people often rely on…

Tutorial videos are a valuable resource for people looking to learn new tasks. People often learn these skills by viewing multiple tutorial videos to get an overall understanding of a task by looking at different approaches to achieve the…

人机交互 · 计算机科学 2025-03-28 Saelyne Yang , Anh Truong , Juho Kim , Dingzeyu Li

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (AR) assistance,…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Lavisha Aggarwal , Vikas Bahirwani , Andrea Colaco

Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people. With the recent progress of vision-language…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Zhiqiang Yuan , Ting Zhang , Ying Deng , Jiapei Zhang , Yeshuang Zhu , Zexi Jia , Jie Zhou , Jinchao Zhang

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Chaoyu Li , Sid Padmanabhuni , Maryam Cheema , Hasti Seifi , Pooyan Fazli

Large-scale multi-task robotic manipulation systems often rely on text to specify the task. In this work, we explore whether a robot can learn by observing humans. To do so, the robot must understand a person's intent and perform the…

While users tend to perceive instructional videos as an experience rather than a lesson with a set of instructions, instructional videos are more effective and appealing than textual user manuals and eliminate the ambiguity in text-based…

人机交互 · 计算机科学 2023-11-22 Songsong Liu , Shu Wang , Kun Sun

Instructors often rely on visual actions such as pointing, marking, and sketching to convey information in educational presentation videos. These subtle visual cues often lack verbal descriptions, forcing low-vision (LV) learners to search…

人机交互 · 计算机科学 2025-08-06 Yotam Sechayk , Ariel Shamir , Amy Pavel , Takeo Igarashi

Goal-oriented planning, or anticipating a series of actions that transition an agent from its current state to a predefined objective, is crucial for developing intelligent assistants aiding users in daily procedural tasks. The problem…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Md Mohaiminul Islam , Tushar Nagarajan , Huiyu Wang , Fu-Jen Chu , Kris Kitani , Gedas Bertasius , Xitong Yang

The reliance on vision for tasks related to cooking and eating healthy can present barriers to cooking for oneself and achieving proper nutrition. There has been little research exploring cooking practices and challenges faced by people…

人机交互 · 计算机科学 2021-07-14 Franklin Mingzhe Li , Jamie Dorst , Peter Cederberg , Patrick Carrington

Over the last decade there has been considerable research into how artificial intelligence (AI), specifically computer vision, can assist people who are blind or have low-vision (BLV) to understand their environment. However, there has been…

人机交互 · 计算机科学 2025-05-27 Bhanuka Gamage , Thanh-Toan Do , Nicholas Seow Chiang Price , Arthur Lowery , Kim Marriott

The rapid growth of virtual reality (VR) has led to increased use of social VR platforms for interaction. However, these platforms lack adequate features to support blind and low vision (BLV) users, posing significant challenges in…

人机交互 · 计算机科学 2026-03-31 Jazmin Collins , Kaylah Myranda Nicholson , Yusuf Khadir , Andrea Stevenson Won , Shiri Azenkot

Locating specific segments within an instructional video is an efficient way to acquire guiding knowledge. Generally, the task of obtaining video segments for both verbal explanations and visual demonstrations is known as visual answer…

计算机视觉与模式识别 · 计算机科学 2025-04-24 Chang Zong , Bin Li , Shoujun Zhou , Jian Wan , Lei Zhang

Good form is the difference between strength and strain, yet for the fast-growing community of at-home fitness enthusiasts, expert feedback is often out of reach. FormCoach transforms a simple camera into an always-on, interactive AI…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Xiaoye Zuo , Nikos Athanasiou , Ginger Delmas , Yiming Huang , Xingyu Fu , Lingjie Liu

Recent advancements in large multimodal models have provided blind or visually impaired (BVI) individuals with new capabilities to interpret and engage with the real world through interactive systems that utilize live video feeds. However,…

人机交互 · 计算机科学 2025-08-06 Ruei-Che Chang , Rosiana Natalie , Wenqian Xu , Jovan Zheng Feng Yap , Anhong Guo

Comparing a user video to a reference how-to video is a key requirement for AR/VR technology delivering personalized assistance tailored to the user's progress. However, current approaches for language-based assistance can only answer…

计算机视觉与模式识别 · 计算机科学 2024-07-01 Tushar Nagarajan , Lorenzo Torresani

We propose L2T, an advancement of visual instruction tuning (VIT). While VIT equips Multimodal LLMs (MLLMs) with promising multimodal capabilities, the current design choices for VIT often result in overfitting and shortcut learning,…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Zhihan Zhou , Feng Hong , Jiaan Luo , Jiangchao Yao , Dongsheng Li , Bo Han , Ya Zhang , Yanfeng Wang

Industrial warehouses are congested with moving forklifts, shelves and personnel, making robot teleoperation particularly risky and demanding for blind and low-vision (BLV) operators. Although accessible teleoperation plays a key role in…

人机交互 · 计算机科学 2025-07-22 Maisha Maimuna , Minhaz Bin Farukee , Sama Nikanfar , Mahfuza Siddiqua , Ayon Roy , Fillia Makedon

We introduce the video detours problem for navigating instructional videos. Given a source video and a natural language query asking to alter the how-to video's current path of execution in a certain way, the goal is to find a related…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Kumar Ashutosh , Zihui Xue , Tushar Nagarajan , Kristen Grauman
‹ 上一页 1 2 3 10 下一页 ›