English
Related papers

Related papers: Vid2Coach: Transforming How-To Videos into Task As…

200 papers

We present Vision-based Navigation with Language-based Assistance (VNLA), a grounded vision-language task where an agent with visual perception is guided via language to find objects in photorealistic indoor environments. The task emulates…

Machine Learning · Computer Science 2019-04-09 Khanh Nguyen , Debadeepta Dey , Chris Brockett , Bill Dolan

Shopping is a routine activity for sighted individuals, yet for people who are blind or have low vision (pBLV), locating and retrieving products in physical environments remains a challenge. This paper presents a multimodal wearable…

Human-Computer Interaction · Computer Science 2026-01-21 Ligao Ruan , Giles Hamilton-Fletcher , Mahya Beheshti , Todd E Hudson , Maurizio Porfiri , John-Ross Rizzo

Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language…

Augmented video presentation tools provide a natural way for presenters to interact with their content, resulting in engaging experiences for remote audiences, such as when a presenter uses hand gestures to manipulate and direct attention…

Human-Computer Interaction · Computer Science 2024-06-27 Temiloluwa Femi-Gege , Matthew Brehmer , Jian Zhao

Large Vision-Language Models (LVLMs) demonstrate a promising direction for assisting individuals with blindness or low-vision (BLV). Yet, measuring their true utility in real-world scenarios is challenging because evaluating whether their…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Eunki Kim , Na Min An , Wan Ju Kang , Sangryul Kim , James Thorne , Hyunjung Shim

Humans are remarkably proficient at controlling their limbs and tools from a wide range of viewpoints and angles, even in the presence of optical distortions. In robotics, this ability is referred to as visual servoing: moving a tool or…

Computer Vision and Pattern Recognition · Computer Science 2017-12-21 Fereshteh Sadeghi , Alexander Toshev , Eric Jang , Sergey Levine

We present V$^2$Dial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conversational tasks. Current multimodal models primarily focus on simpler tasks (e.g., VQA,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Adnen Abdessaied , Anna Rohrbach , Marcus Rohrbach , Andreas Bulling

Hybrid tutoring, where a human tutor supports multiple students in learning with educational technology, is an increasingly common application to deliver high-impact tutoring at scale. However, past hybrid tutoring applications are limited…

Human-Computer Interaction · Computer Science 2025-05-14 Eason Chen , Xinyi Tang , Aprille Xi , Chenyu Lin , Conrad Borchers , Shivang Gupta , Jionghao Lin , Kenneth R Koedinger

Vision-Language Models (VLMs) acquire real-world knowledge and general reasoning ability through Internet-scale image-text corpora. They can augment robotic systems with scene understanding and task planning, and assist visuomotor policies…

Robotics · Computer Science 2025-06-23 Kaiyuan Chen , Shuangyu Xie , Zehan Ma , Pannag R Sanketi , Ken Goldberg

Text-guided image-to-video (I2V) generation aims to generate a coherent video that preserves the identity of the input image and semantically aligns with the input prompt. Existing methods typically augment pretrained text-to-video (T2V)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-28 Xun Guo , Mingwu Zheng , Liang Hou , Yuan Gao , Yufan Deng , Pengfei Wan , Di Zhang , Yufan Liu , Weiming Hu , Zhengjun Zha , Haibin Huang , Chongyang Ma

Large Vision-Language Models (VLMs) excel at understanding and generating video descriptions but their high memory, computation, and deployment demands hinder practical use particularly for blind and low-vision (BLV) users who depend on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Shruti Singh Baghel , Yash Pratap Singh Rathore , Sushovan Jena , Anurag Pradhan , Amit Shukla , Arnav Bhavsar , Pawan Goyal

Robotic guidance systems have shown promise in supporting blind and visually impaired (BVI) individuals with wayfinding and obstacle avoidance. However, most existing systems assume a clear path and do not support a critical aspect of…

Robotics · Computer Science 2026-03-17 Shaojun Cai , Nuwan Janaka , Ashwin Ram , Janidu Shehan , Yingjia Wan , Kotaro Hara , David Hsu

Cooking is a vital yet challenging activity for people with visual impairments (PVI). It involves tasks that can be dangerous or difficult without vision, such as handling a knife or adding a suitable amount of salt. A better understanding…

Human-Computer Interaction · Computer Science 2023-10-10 Ru Wang , Nihan Zhou , Tam Nguyen , Sanbrita Mondal , Bilge Mutlu , Yuhang Zhao

Long video understanding is a significant and ongoing challenge in the intersection of multimedia and artificial intelligence. Employing large language models (LLMs) for comprehending video becomes an emerging and promising method. However,…

Computation and Language · Computer Science 2024-08-27 Yunxin Li , Xinyu Chen , Baotain Hu , Min Zhang

We investigate whether tactile charts support comprehension and learning of complex visualizations for blind and low-vision (BLV) individuals and contribute four tactile chart designs and an interview study. Visualizations are powerful…

Human-Computer Interaction · Computer Science 2025-08-11 Tingying He , Maggie McCracken , Daniel Hajas , Sarah Creem-Regehr , Alexander Lex

Many decision-making tasks, where both accuracy and efficiency matter, still require human supervision. For example, tasks like traffic officers reviewing hour-long dashcam footage or researchers screening conference videos can benefit from…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Shenghui Chen , Po-han Li , Sandeep Chinchali , Ufuk Topcu

We introduce VisIT-Bench (Visual InsTruction Benchmark), a benchmark for evaluation of instruction-following vision-language models for real-world use. Our starting point is curating 70 'instruction families' that we envision instruction…

Computation and Language · Computer Science 2023-12-27 Yonatan Bitton , Hritik Bansal , Jack Hessel , Rulin Shao , Wanrong Zhu , Anas Awadalla , Josh Gardner , Rohan Taori , Ludwig Schmidt

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Shenghao Fu , Qize Yang , Yuan-Ming Li , Yi-Xing Peng , Kun-Yu Lin , Xihan Wei , Jian-Fang Hu , Xiaohua Xie , Wei-Shi Zheng

Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy…

Robotics · Computer Science 2025-12-22 Jonas Pai , Liam Achenbach , Victoriano Montesinos , Benedek Forrai , Oier Mees , Elvis Nava

Visual planning, by offering a sequence of intermediate visual subgoals to a goal-conditioned low-level policy, achieves promising performance on long-horizon manipulation tasks. To obtain the subgoals, existing methods typically resort to…

Robotics · Computer Science 2025-08-08 Wenyan Yang , Ahmet Tikna , Yi Zhao , Yuying Zhang , Luigi Palopoli , Marco Roveri , Joni Pajarinen