English
Related papers

Related papers: Can ChatGPT assist visually impaired people with m…

200 papers

Recent advancements in large multimodal models have provided blind or visually impaired (BVI) individuals with new capabilities to interpret and engage with the real world through interactive systems that utilize live video feeds. However,…

Human-Computer Interaction · Computer Science 2025-08-06 Ruei-Che Chang , Rosiana Natalie , Wenqian Xu , Jovan Zheng Feng Yap , Anhong Guo

Artificial intelligence-based chatbots are increasingly influencing physics education due to their ability to interpret and respond to textual and visual inputs. This study evaluates the performance of two large multimodal model-based…

Physics Education · Physics 2025-05-29 Giulia Polverini , Jakob Melin , Elias Onerud , Bor Gregorcic

Indoor navigation presents unique challenges due to complex layouts and the unavailability of GNSS signals. Existing solutions often struggle with contextual adaptation, and typically require dedicated hardware. In this work, we explore the…

Artificial Intelligence · Computer Science 2025-06-23 Alberto Coffrini , Paolo Barsocchi , Francesco Furfari , Antonino Crivello , Alessio Ferrari

Assistive technologies for people with visual impairments (PVI) have made significant advancements, particularly with the integration of artificial intelligence (AI) and real-time sensor technologies. However, current solutions often…

Human-Computer Interaction · Computer Science 2024-10-08 He Zhang , Nicholas J. Falletta , Jingyi Xie , Rui Yu , Sooyeon Lee , Syed Masum Billah , John M. Carroll

Despite the significant advancements in natural language processing capabilities demonstrated by large language models such as ChatGPT, their proficiency in comprehending and processing spatial information, especially within the domains of…

Computation and Language · Computer Science 2023-12-07 He Yan , Xinyao Hu , Xiangpeng Wan , Chengyu Huang , Kai Zou , Shiqi Xu

Vision is essential for human navigation. The World Health Organization (WHO) estimates that 43.3 million people were blind in 2020, and this number is projected to reach 61 million by 2050. Modern scene understanding models could empower…

While there is no replacement for the learned expertise, devotion, and social benefits of a guide dog, there are cases in which a robot navigation assistant could be helpful for individuals with blindness or low vision (BLV). This study…

Robotics · Computer Science 2024-06-10 Rayna Hata , Narit Trikasemsak , Andrea Giudice , Stacy A. Doore

Objective: To evaluate the efficiency of large language models (LLMs) such as ChatGPT to assist in diagnosing neuro-ophthalmic diseases based on detailed case descriptions. Methods: We selected 22 different case reports of neuro-ophthalmic…

Computers and Society · Computer Science 2023-09-25 Yeganeh Madadi , Mohammad Delsoz , Priscilla A. Lao , Joseph W. Fong , TJ Hollingsworth , Malik Y. Kahook , Siamak Yousefi

The well-known artificial intelligence-based chatbot ChatGPT-4 has become able to process image data as input in October 2023. We investigated its performance on the Test of Understanding Graphs in Kinematics to inform the physics education…

Physics Education · Physics 2025-02-10 Giulia Polverini , Bor Gregorcic

Ensuring accessible pedestrian navigation requires reasoning about both semantic and spatial aspects of complex urban scenes, a challenge that existing Large Vision-Language Models (LVLMs) struggle to meet. Although these models can…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Rafi Ibn Sultan , Hui Zhu , Xiangyu Zhou , Chengyin Li , Prashant Khanduri , Marco Brocanelli , Dongxiao Zhu

LLM-Glasses is a wearable navigation system which assists visually impaired people by utilizing YOLO-World object detection, GPT-4o-based reasoning, and haptic feedback for real-time guidance. The device translates visual scene…

Human-Computer Interaction · Computer Science 2026-01-21 Issatay Tokmurziyev , Miguel Altamirano Cabrera , Muhammad Haris Khan , Yara Mahmoud , Dzmitry Tsetserukou

This study investigates the potential of a multimodal large language model (LLM), specifically ChatGPT-4o, to perform human-like interpretations of traffic scenes using static dashcam images. Herein, we focus on three judgment tasks…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Yuki Yoshihara , Linjing Jiang , Nihan Karatas , Hitoshi Kanamori , Asuka Harada , Takahiro Tanaka

Studies on in-vehicle conversational agents have traditionally relied on pre-scripted prompts or limited voice commands, constraining natural driver-agent interaction. To resolve this issue, the present study explored the potential of a…

Human-Computer Interaction · Computer Science 2025-08-12 Yeana Lee Bond , Mungyeong Choe , Baker Kasim Hasan , Arsh Siddiqui , Myounghoon Jeon

Visual Language Navigation (VLN) powered robots have the potential to guide blind people by understanding route instructions provided by sighted passersby. This capability allows robots to operate in environments often unknown a prior.…

Robotics · Computer Science 2026-01-29 Masaki Kuribayashi , Kohei Uehara , Allan Wang , Daisuke Sato , Simon Chu , Shigeo Morishima

Data visualization creators often lack formal training, resulting in a knowledge gap in design practice. Large language models such as ChatGPT, with their vast internet-scale training data, offer transformative potential to address this…

Human-Computer Interaction · Computer Science 2025-05-20 Nam Wook Kim , Yongsu Ahn , Grace Myers , Benjamin Bach

According to the World Health Organization, visual impairment is estimated to affect approximately 2.2 billion people worldwide. The visually impaired must currently rely on navigational aids to replace their sense of sight, like a white…

Human-Computer Interaction · Computer Science 2022-06-23 Stanley Shen

Web accessibility ensures that individuals with disabilities can access and interact with digital content without barriers, yet a significant majority of most used websites fail to meet accessibility standards. This study evaluates…

Human-Computer Interaction · Computer Science 2025-07-21 Ammar Ahmed , Margarida Fresco , Fredrik Forsberg , Hallvard Grotli

Recently, the flourishing large language models(LLM), especially ChatGPT, have shown exceptional performance in language understanding, reasoning, and interaction, attracting users and researchers from multiple fields and domains. Although…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Haonan Guo , Xin Su , Chen Wu , Bo Du , Liangpei Zhang , Deren Li

We present an interactive visual framework named InternGPT, or iGPT for short. The framework integrates chatbots that have planning and reasoning capabilities, such as ChatGPT, with non-verbal instructions like pointing movements that…

Computer Vision and Pattern Recognition · Computer Science 2023-06-05 Zhaoyang Liu , Yinan He , Wenhai Wang , Weiyun Wang , Yi Wang , Shoufa Chen , Qinglong Zhang , Zeqiang Lai , Yang Yang , Qingyun Li , Jiashuo Yu , Kunchang Li , Zhe Chen , Xue Yang , Xizhou Zhu , Yali Wang , Limin Wang , Ping Luo , Jifeng Dai , Yu Qiao

The use of AI assistants, along with the challenges they present, has sparked significant debate within the community of computer science education. While these tools demonstrate the potential to support students' learning and instructors'…

Human-Computer Interaction · Computer Science 2023-11-14 Tianjia Wang , Daniel Vargas-Díaz , Chris Brown , Yan Chen
‹ Prev 1 2 3 10 Next ›