中文
相关论文

相关论文: Explore, Listen, Inspect: Supporting Multimodal In…

200 篇论文

Online shopping has become a valuable modern convenience, but blind or low vision (BLV) users still face significant challenges using it, because of: 1) inadequate image descriptions and 2) the inability to filter large amounts of…

人机交互 · 计算机科学 2021-02-02 Ruolin Wang , Zixuan Chen , Mingrui "Ray" Zhang , Zhaoheng Li , Zhixiu Liu , Zihan Dang , Chun Yu , Xiang "Anthony" Chen

Multimodal large language models (MLLMs) are changing how Blind and Low Vision (BLV) people access visual information. Unlike traditional visual interpretation tools that only provide descriptions, MLLM-enabled applications offer…

人机交互 · 计算机科学 2026-02-20 Ricardo E. Gonzalez Penuela , Crescentia Jung , Sharon Y Lin , Ruiying Hu , Shiri Azenkot

Robust humanoid locomotion requires accurate and globally consistent perception of the surrounding 3D environment. However, existing perception modules, mainly based on depth images or elevation maps, offer only partial and locally…

机器人学 · 计算机科学 2025-11-19 Qingwei Ben , Botian Xu , Kailin Li , Feiyu Jia , Wentao Zhang , Jingping Wang , Jingbo Wang , Dahua Lin , Jiangmiao Pang

People with blindness and low vision (pBLV) face significant challenges, struggling to navigate environments and locate objects due to limited visual cues. Spatial reasoning is crucial for these individuals, as it enables them to understand…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Alexey Magay , Dhurba Tripathi , Yu Hao , Yi Fang

Integrating LiDAR and camera information in the bird's eye view (BEV) representation has demonstrated its effectiveness in 3D object detection. However, because of the fundamental disparity in geometric accuracy between these sensors,…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Guowen Zhang , Chenhang He , Liyi Chen , Lei Zhang

Although recent efforts have developed accessible data visualization tools for blind and low-vision (BLV) users, most follow a "design for them" approach that creates an unintentional divide between sighted creators and BLV consumers. This…

人机交互 · 计算机科学 2025-09-18 JooYoung Seo , Saairam Venkatesh , Daksh Pokar , Sanchita Kamath , Krishna Anandan Ganesan

Blind and low vision (BLV) users often rely on alt text to understand what a digital image is showing. However, recent research has investigated how touch-based image exploration on touchscreens can supplement alt text. Touchscreen-based…

人机交互 · 计算机科学 2023-02-21 Vishnu Nair , Hanxiu 'Hazel' Zhu , Brian A. Smith

360{\deg} videos enable users to freely choose their viewing paths, but blind and low vision (BLV) users are often excluded from this interactive experience. To bridge this gap, we present Branch Explorer, a system that transforms 360{\deg}…

人机交互 · 计算机科学 2025-07-15 Shuchang Xu , Xiaofu Jin , Wenshuo Zhang , Huamin Qu , Yukang Yan

Gaze interaction presents a promising avenue in Virtual Reality (VR) due to its intuitive and efficient user experience. Yet, the depth control inherent in our visual system remains underutilized in current methods. In this study, we…

人机交互 · 计算机科学 2024-05-08 Chenyang Zhang , Tiansu Chen , Eric Shaffer , Elahe Soltanaghai

Traditional accessibility methods like alternative text and data tables typically underrepresent data visualization's full potential. Keyboard-based chart navigation has emerged as a potential solution, yet efficient data exploration…

人机交互 · 计算机科学 2024-08-20 Joshua Gorniak , Yoon Kim , Donglai Wei , Nam Wook Kim

While audio description (AD) is the standard approach for making videos accessible to blind and low vision (BLV) people, existing AD guidelines do not consider BLV users' varied preferences across viewing scenarios. These scenarios range…

人机交互 · 计算机科学 2024-03-19 Lucy Jiang , Crescentia Jung , Mahika Phutane , Abigale Stangl , Shiri Azenkot

Visualization dashboards are regularly used for data exploration and analysis, but their complex interactions and interlinked views often require time-consuming onboarding sessions from dashboard authors. Preparing these onboarding…

人机交互 · 计算机科学 2026-03-04 Vaishali Dhanoa , Gabriela Molina León , Eve Hoggan , Eduard Gröller , Marc Streit , Niklas Elmqvist

Recent advances in multimodal recommendation (MMR) highlight the potential of integrating visual and textual content to enrich item representations. However, existing methods often rely on coarse visual features and naive fusion strategies,…

信息检索 · 计算机科学 2025-11-11 Hai-Dang Kieu , Min Xu , Thanh Trung Huynh , Dung D. Le

Virtual Reality (VR) enables users to collaborate while exploring scenarios not realizable in the physical world. We propose CollabVR, a distributed multi-user collaboration environment, to explore how digital content improves expression…

人机交互 · 计算机科学 2019-11-19 Zhenyi He , Karl Rosenberg , Ken Perlin

People with blindness and low vision (pBLV) experience significant challenges when locating final destinations or targeting specific objects in unfamiliar environments. Furthermore, besides initially locating and orienting oneself to a…

计算机视觉与模式识别 · 计算机科学 2022-08-19 Yu Hao , Junchi Feng , John-Ross Rizzo , Yao Wang , Yi Fang

3D object detection from multiple image views is a fundamental and challenging task for visual scene understanding. Owing to its low cost and high efficiency, multi-view 3D object detection has demonstrated promising application prospects.…

计算机视觉与模式识别 · 计算机科学 2022-11-18 Zehui Chen , Zhenyu Li , Shiquan Zhang , Liangji Fang , Qinhong Jiang , Feng Zhao

In this paper, we design a multimodal framework for object detection, recognition and mapping based on the fusion of stereo camera frames, point cloud Velodyne Lidar scans, and Vehicle-to-Vehicle (V2V) Basic Safety Messages (BSMs) exchanged…

计算机视觉与模式识别 · 计算机科学 2017-05-25 Yassine Maalej , Sameh Sorour , Ahmed Abdel-Rahim , Mohsen Guizani

Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scene understanding. This deficiency manifests as three intertwined challenges: weak…

机器人学 · 计算机科学 2026-05-29 Zhongyu Xia , Yousen Tang , Bingqing Wei , Yongtao Wang

Video-based learning (VBL) has become a dominant method for learning practical skills, yet accessibility guidelines provide limited guidance for users with cognitive differences. In particular, challenges that individuals with Borderline…

人机交互 · 计算机科学 2026-02-10 Hyehyun Chu , Seungju Kim , Chen Zhou , Yu-Kai Hung , Saelyne Yang , Hyun W. Ka , Juho Kim

This paper explores the effectiveness of Multimodal Large Language models (MLLMs) as assistive technologies for visually impaired individuals. We conduct a user survey to identify adoption patterns and key challenges users face with such…