English
Related papers

Related papers: Vid2Coach: Transforming How-To Videos into Task As…

200 papers

Distilling knowledge from human demonstrations is a promising way for robots to learn and act. Existing methods, which often rely on coarsely-aligned video pairs, are typically constrained to learning global or task-level features. As a…

Robotics · Computer Science 2025-11-18 Sicheng Xie , Haidong Cao , Zejia Weng , Zhen Xing , Haoran Chen , Shiwei Shen , Jiaqi Leng , Zuxuan Wu , Yu-Gang Jiang

Robots can use Visual Imitation Learning (VIL) to learn manipulation tasks from video demonstrations. However, translating visual observations into actionable robot policies is challenging due to the high-dimensional nature of video data.…

Robotics · Computer Science 2025-01-22 Ananth Jonnavittula , Sagar Parekh , Dylan P. Losey

Tutorial videos of mobile apps have become a popular and compelling way for users to learn unfamiliar app features. To make the video accessible to the users, video creators always need to annotate the actions in the video, including what…

Human-Computer Interaction · Computer Science 2023-08-08 Sidong Feng , Chunyang Chen , Zhenchang Xing

Cooking requires not only following instructions but also understanding, executing, and monitoring each step - a process that can be challenging without visual guidance. Although recipe images and videos offer helpful cues, they often lack…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Oleh Kuzyk , Zuoyue Li , Marc Pollefeys , Xi Wang

This paper presents a VR-based guide dog training system designed to assist novice trainers in understanding guide dog behavior and issuing appropriate training commands. Guide dogs play a vital role in supporting independent mobility for…

Human-Computer Interaction · Computer Science 2025-10-29 Qirong Zhu , Ansheng Wang , Shinji Tanaka , Yasutoshi Makino , Hiroyuki Shinoda

While image editing has advanced rapidly, video editing remains less explored, facing challenges in consistency, control, and generalization. We study the design space of data, architecture, and control, and introduce \emph{EasyV2V}, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jinjie Mai , Chaoyang Wang , Guocheng Gordon Qian , Willi Menapace , Sergey Tulyakov , Bernard Ghanem , Peter Wonka , Ashkan Mirzaei

Multimodal large language models (MLLMs) are changing how Blind and Low Vision (BLV) people access visual information. Unlike traditional visual interpretation tools that only provide descriptions, MLLM-enabled applications offer…

Human-Computer Interaction · Computer Science 2026-02-20 Ricardo E. Gonzalez Penuela , Crescentia Jung , Sharon Y Lin , Ruiying Hu , Shiri Azenkot

"Scene description" applications that describe visual content in a photo are useful daily tools for blind and low vision (BLV) people. Researchers have studied their use, but they have only explored those that leverage remote sighted…

Human-Computer Interaction · Computer Science 2025-03-13 Ricardo Gonzalez , Jazmin Collins , Shiri Azenkot , Cynthia Bennett

We present BioCoach, a biomechanics-grounded vision--language framework for fitness coaching from streaming video. BioCoach fuses visual appearance and 3D skeletal kinematics, through a novel three-stage pipeline: an exercise-specific…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yuyang Ji , Yixuan Shen , Shengjie Zhu , Yu Kong , Feng Liu

Vision plays a crucial role in comprehending the world around us. More than 85% of the external information is obtained through the vision system. It influences our mobility, cognition, information access, and interaction with the…

Computer Vision and Pattern Recognition · Computer Science 2022-01-05 Ishwarya Sivakumar , Nishaali Meenakshisundaram , Ishwarya Ramesh , Shiloah Elizabeth D , Sunil Retmin Raj C

Video content remains largely inaccessible to blind and low-vision (BLV) users. To address this, we introduce a prototype that leverages a multimodal agent - powered by a novel conversational architecture using a multimodal large language…

Video summarization aims to create short, accurate, and cohesive summaries of longer videos. Despite the existence of various video summarization datasets, a notable limitation is their limited amount of source videos, which hampers the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Hang Hua , Yolo Yunlong Tang , Chenliang Xu , Jiebo Luo

Video Internet of Things (VIoT) has shown full potential in collecting an unprecedented volume of video data. How to schedule the domain-specific perceiving models and analyze the collected videos uniformly, efficiently, and especially…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Yaoyao Zhong , Mengshi Qi , Rui Wang , Yuhan Qiu , Yang Zhang , Huadong Ma

Visually Impaired Assistance (VIA) aims to automatically help the visually impaired (VI) handle daily activities. The advancement of VIA primarily depends on developments in Computer Vision (CV) and Natural Language Processing (NLP), both…

Computation and Language · Computer Science 2024-02-13 Yi Zhao , Yilin Zhang , Rong Xiang , Jing Li , Hillming Li

The advent and proliferation of large multi-modal models (LMMs) have introduced new paradigms to computer vision, transforming various tasks into a unified visual question answering framework. Video Quality Assessment (VQA), a classic field…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Ziheng Jia , Zicheng Zhang , Jiaying Qian , Haoning Wu , Wei Sun , Chunyi Li , Xiaohong Liu , Weisi Lin , Guangtao Zhai , Xiongkuo Min

Videos make exercise instruction widely available, but they rely on visual demonstrations that blind and low vision (BLV) learners cannot see. While audio descriptions (AD) can make videos accessible, describing movements remains…

Human-Computer Interaction · Computer Science 2026-03-02 Ujjaini Das , Shreya Kappala , Meng Chen , Mina Huh , Amy Pavel

The advances in AI-enabled techniques have accelerated the creation and automation of visualizations in the past decade. However, presenting visualizations in a descriptive and generative format remains a challenge. Moreover, current…

Human-Computer Interaction · Computer Science 2024-03-28 Qing Chen , Ying Chen , Ruishi Zou , Wei Shuai , Yi Guo , Jiazhe Wang , Nan Cao

While there is no replacement for the learned expertise, devotion, and social benefits of a guide dog, there are cases in which a robot navigation assistant could be helpful for individuals with blindness or low vision (BLV). This study…

Robotics · Computer Science 2024-06-10 Rayna Hata , Narit Trikasemsak , Andrea Giudice , Stacy A. Doore

Alongside the prevalence of mobile videos, the general public leans towards consuming vertical videos on hand-held devices. To revitalize the exposure of horizontal contents, we hereby set forth the exploration of automated…

Computer Vision and Pattern Recognition · Computer Science 2021-06-24 Tun Zhu , Daoxin Zhang , Yao Hu , Tianran Wang , Xiaolong Jiang , Jianke Zhu , Jiawei Li

Modern manufacturing demands robotic assembly systems with enhanced flexibility and reliability. However, traditional approaches often rely on programming tailored to each product by experts for fixed settings, which are inherently…

Robotics · Computer Science 2025-09-23 Xiwei Zhao , Yiwei Wang , Yansong Wu , Fan Wu , Teng Sun , Zhonghua Miao , Sami Haddadin , Alois Knoll