English
Related papers

Related papers: mmWalk: Towards Multi-modal Multi-view Walking Ass…

200 papers

Industrial warehouses are congested with moving forklifts, shelves and personnel, making robot teleoperation particularly risky and demanding for blind and low-vision (BLV) operators. Although accessible teleoperation plays a key role in…

Human-Computer Interaction · Computer Science 2025-07-22 Maisha Maimuna , Minhaz Bin Farukee , Sama Nikanfar , Mahfuza Siddiqua , Ayon Roy , Fillia Makedon

People with blindness and low vision (pBLV) experience significant challenges when locating final destinations or targeting specific objects in unfamiliar environments. Furthermore, besides initially locating and orienting oneself to a…

Computer Vision and Pattern Recognition · Computer Science 2022-08-19 Yu Hao , Junchi Feng , John-Ross Rizzo , Yao Wang , Yi Fang

In recent decades, several assistive technologies have been developed to improve the ability of blind and visually impaired (BVI) individuals to navigate independently and safely. At the same time, simultaneous localization and mapping…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Marziyeh Bamdad , Davide Scaramuzza , Alireza Darvishy

This paper investigates the potential of vision-language models (VLMs) to assist people with blindness and low vision (pBLV) in navigation tasks. We evaluate state-of-the-art closed-source models, including GPT-4V, GPT-4o, Gemini-1.5-Pro,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Yu Li , Yuchen Zheng , Giles Hamilton-Fletcher , Marco Mezzavilla , Yao Wang , Sundeep Rangan , Maurizio Porfiri , Zhou Yu , John-Ross Rizzo

Large Vision-Language Models (LVLMs) demonstrate a promising direction for assisting individuals with blindness or low-vision (BLV). Yet, measuring their true utility in real-world scenarios is challenging because evaluating whether their…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Eunki Kim , Na Min An , Wan Ju Kang , Sangryul Kim , James Thorne , Hyunjung Shim

This paper presents a wearable assistive device with the shape of a pair of eyeglasses that allows visually impaired people to navigate safely and quickly in unfamiliar environment, as well as perceive the complicated environment to…

Computer Vision and Pattern Recognition · Computer Science 2019-06-25 Jinqiang Bai , Zhaoxiang Liu , Yimin Lin , Ye Li , Shiguo Lian , Dijun Liu

For individuals with blindness or low vision (BLV), navigating complex environments can pose serious risks. Large Vision-Language Models (LVLMs) show promise for generating scene descriptions, but their effectiveness for BLV users remains…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Na Min An , Eunki Kim , Wan Ju Kang , Sangryul Kim , James Thorne , Hyunjung Shim

Actively participating in body movement such as dance, sports, and fitness activities is challenging for people who are blind or have low vision (BLV). Teachers primarily rely on verbal instructions and physical demonstrations with limited…

Human-Computer Interaction · Computer Science 2024-10-01 Madhuka Thisuri De Silva , Sarah Goodwin , Leona M Holloway , Matthew Butler

By the year 2023, the global population of individuals with impaired vision has surpassed 220 million. People with impaired vision will find it difficult while finding path or avoiding obstacles, and must ask for auxiliary tools for help.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Feifan Yan , Tianle Zeng , Meixi He

The visual world around us constantly evolves, from real-time news and social media trends to global infrastructure changes visible through satellite imagery and augmented reality enhancements. However, Multimodal Large Language Models…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Mingyang Fu , Yuyang Peng , Dongping Chen , Zetong Zhou , Benlin Liu , Yao Wan , Zhou Zhao , Philip S. Yu , Ranjay Krishna

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in interpreting visual layouts and text. However, a significant challenge remains in their ability to interpret robustly and reason over multi-tabular data presented as…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Anshul Singh , Chris Biemann , Jan Strich

The development of smart cities has led to the generation of massive amounts of multi-modal data in the context of a range of tasks that enable a comprehensive monitoring of the smart city infrastructure and services. This paper surveys one…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Zhangyong Tang , Tianyang Xu , Xuefeng Zhu , Hui Li , Shaochuan Zhao , Tao Zhou , Chunyang Cheng , Xiaojun Wu , Josef Kittler

Multimodal large language models (MLLMs) are changing how Blind and Low Vision (BLV) people access visual information. Unlike traditional visual interpretation tools that only provide descriptions, MLLM-enabled applications offer…

Human-Computer Interaction · Computer Science 2026-02-20 Ricardo E. Gonzalez Penuela , Crescentia Jung , Sharon Y Lin , Ruiying Hu , Shiri Azenkot

We introduce WearVQA, the first benchmark specifically designed to evaluate the Visual Question Answering (VQA) capabilities of multi-model AI assistant on wearable devices like smart glasses. Unlike prior benchmarks that focus on…

To help the blind people walk to the destination efficiently and safely in indoor environment, a novel wearable navigation device is presented in this paper. The locating, way-finding, route following and obstacle avoiding modules are the…

Computer Vision and Pattern Recognition · Computer Science 2019-05-01 Jinqiang Bai , Shiguo Lian , Zhaoxiang Liu , Kai Wang , Dijun Liu

Passive tracking methods, such as phone and wearable sensing, have become dominant in monitoring human behaviors in modern ubiquitous computing studies. While there have been significant advances in machine-learning approaches to translate…

Human-Computer Interaction · Computer Science 2025-10-31 Jiachen Li , Xiwen Li , Justin Steinberg , Akshat Choube , Bingsheng Yao , Xuhai Xu , Dakuo Wang , Elizabeth Mynatt , Varun Mishra

Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. However, ensuring the safety of these models remains a significant challenge,…

Computation and Language · Computer Science 2025-06-04 Wenxuan Wang , Xiaoyuan Liu , Kuiyi Gao , Jen-tse Huang , Youliang Yuan , Pinjia He , Shuai Wang , Zhaopeng Tu

Blind and Low Vision (BLV) people have adopted AI-powered visual interpretation applications to address their daily needs. While these applications have been helpful, prior work has found that users remain unsatisfied by their frequent…

Human-Computer Interaction · Computer Science 2025-03-11 Ricardo E. Gonzalez Penuela , Ruiying Hu , Sharon Lin , Tanisha Shende , Shiri Azenkot

With the proliferation of low altitude unmanned aerial vehicles (UAVs), visual multi-object tracking is becoming a critical security technology, demanding significant robustness even in complex environmental conditions. However, tracking…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Tianyang Xu , Jinjie Gu , Xuefeng Zhu , XiaoJun Wu , Josef Kittler

Effective visual accessibility in Virtual Reality (VR) is crucial for Blind and Low Vision (BLV) users. However, designing visual accessibility systems is challenging due to the complexity of 3D VR environments and the need for techniques…

Human-Computer Interaction · Computer Science 2025-02-07 Junlong Chen , Rosella P. Galindo Esparza , Vanja Garaj , Per Ola Kristensson , John Dudley