English
Related papers

Related papers: Making Videos Accessible for Blind and Low Vision …

200 papers

Effective time management during presentations is challenging, particularly for Blind and Low-Vision (BLV) individuals, as existing tools often lack accessibility and multimodal feedback. To address this gap, we developed vashTimer: a free,…

Human-Computer Interaction · Computer Science 2025-09-25 Aziz N Zeidieh , Sanchita S. Kamath , JooYoung Seo

Multimodal Large Language Models (MLLMs) hold immense promise as assistive technologies for the blind and visually impaired (BVI) community. However, we identify a critical failure mode that undermines their trustworthiness in real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xiantao Zhang

Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily lives. However, these models often produce errors that are difficult to detect without…

Human-Computer Interaction · Computer Science 2025-07-22 Meng Chen , Akhil Iyer , Amy Pavel

Shopping is a routine activity for sighted individuals, yet for people who are blind or have low vision (pBLV), locating and retrieving products in physical environments remains a challenge. This paper presents a multimodal wearable…

Human-Computer Interaction · Computer Science 2026-01-21 Ligao Ruan , Giles Hamilton-Fletcher , Mahya Beheshti , Todd E Hudson , Maurizio Porfiri , John-Ross Rizzo

Access to non-verbal cues in social interactions is vital for people with visual impairment. It has been shown that non-verbal cues such as eye contact, number of people, their names and positions are helpful for individuals who are blind.…

Computers and Society · Computer Science 2017-11-30 M. Saquib Sarfraz , Angela Constantinescu , Melanie Zuzej , Rainer Stiefelhagen

Audio description (AD) makes video content accessible to blind and low-vision (BLV) audiences, but producing high-quality descriptions is resource-intensive. Automated AD offers scalability, and prior studies show human-in-the-loop editing…

Human-Computer Interaction · Computer Science 2026-02-04 Lana Do , Shasta Ihorn , Charity Pitcher-Cooper , Juvenal Francisco Barajas , Gio Jung , Xuan Duy Anh Nguyen , Sanjay Mirani , Ilmi Yoon

Combining conversational AI with refreshable tactile displays (RTDs) offers significant potential for creating accessible data visualization for people who are blind or have low vision (BLV). To support researchers and developers building…

Human-Computer Interaction · Computer Science 2026-02-18 Samuel Reinders , Munazza Zaib , Matthew Butler , Bongshin Lee , Ingrid Zukerman , Lizhen Qu , Kim Marriott

Current scene perception tools for Blind and Low Vision (BLV) individuals rely on spoken descriptions but lack engaging representations of visually pleasing distant environmental landscapes (Vista spaces). Our proposed Scene2Audio framework…

Human-Computer Interaction · Computer Science 2026-03-31 Chitralekha Gupta , Jing Peng , Ashwin Ram , Shreyas Sridhar , Christophe Jouffrais , Suranga Nanayakkara

Despite recent advances, long-sequence video generation frameworks still suffer from significant limitations: poor assistive capability, suboptimal visual quality, and limited expressiveness. To mitigate these limitations, we propose MAViS,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Qian Wang , Ziqi Huang , Ruoxi Jia , Paul Debevec , Ning Yu

User interface (UI) agents promise to make inaccessible or complex UIs easier to access for blind and low-vision (BLV) users. However, current UI agents typically perform tasks end-to-end without involving users in critical choices or…

Human-Computer Interaction · Computer Science 2025-09-01 Yi-Hao Peng , Dingzeyu Li , Jeffrey P. Bigham , Amy Pavel

Existing MLLMs encounter significant challenges in modeling the temporal context within long videos. Currently, mainstream Agent-based methods use external tools to assist a single MLLM in answering long video questions. Despite such…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Boyu Chen , Zhengrong Yue , Siran Chen , Zikang Wang , Yang Liu , Peng Li , Yali Wang

The proliferation of mobile devices and social media has revolutionized content dissemination, with short-form video becoming increasingly prevalent. This shift has introduced the challenge of video reframing to fit various screen aspect…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Jiawang Cao , Yongliang Wu , Weiheng Chi , Wenbo Zhu , Ziyue Su , Jay Wu

Large Vision-Language Models (LVLMs) empower autonomous mobile agents, yet their security under realistic mobile deployment constraints remains underexplored. While agents are vulnerable to visual prompt injections, stealthily executing…

Cryptography and Security · Computer Science 2026-04-10 Renhua Ding , Xiao Yang , Zhengwei Fang , Jun Luo , Kun He , Jun Zhu

Navigating unfamiliar environments remains one of the most persistent and critical challenges for people who are blind or have limited vision (BLV). Existing assistive tools often rely on online services or APIs, making them costly,…

Human-Computer Interaction · Computer Science 2025-10-28 Dabbrata Das , Argho Deb Das , Farhan Sadaf , Azhar Uddin , Tirtho Mondal

While there is no replacement for the learned expertise, devotion, and social benefits of a guide dog, there are cases in which a robot navigation assistant could be helpful for individuals with blindness or low vision (BLV). This study…

Robotics · Computer Science 2024-06-10 Rayna Hata , Narit Trikasemsak , Andrea Giudice , Stacy A. Doore

Recent advances in Large Language Models (LLMs) have significantly improved natural language understanding and generation, enhancing Human-Computer Interaction (HCI). However, LLMs are limited to unimodal text processing and lack the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Chenxi Li

With the advancement of multimodal large language models (MLLMs) and coding agents, the website development has shifted from manual programming to agent-based project-level code synthesis. Existing benchmarks rely on idealized assumptions,…

Artificial Intelligence · Computer Science 2026-05-01 Qiyao Wang , Haoran Hu , Longze Chen , Hongbo Wang , Hamid Alinejad-Rokny , Yuan Lin , Min Yang

Video accessibility is crucial for blind and low vision users for equitable engagements in education, employment, and entertainment. Despite the availability of professional and amateur services and tools, most human-generated descriptions…

Human-Computer Interaction · Computer Science 2022-01-12 Shasta Ihorn , Yue-Ting Siu , Aditya Bodi , Lothar Narins , Jose M. Castanon , Yash Kant , Abhishek Das , Ilmi Yoon , Pooyan Fazli

People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our observations of visual rehabilitation therapists (VRTs)…

Human-Computer Interaction · Computer Science 2025-07-28 Mina Huh , Zihui Xue , Ujjaini Das , Kumar Ashutosh , Kristen Grauman , Amy Pavel

Domestic AI agents faces ethical, autonomy, and inclusion challenges, particularly for overlooked groups like children, elderly, and Neurodivergent users. We present the Plural Voices Model (PVM), a novel single-agent framework that…

Human-Computer Interaction · Computer Science 2025-10-23 Joydeep Chandra , Satyam Kumar Navneet
‹ Prev 1 3 4 5 6 7 10 Next ›