中文
相关论文

相关论文: Tools and Tasks in Sensemaking: A Visual Accessibi…

200 篇论文

Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in Vision-Language Models (VLMs), achieving human-level VSR…

This paper investigates the potential of vision-language models (VLMs) to assist people with blindness and low vision (pBLV) in navigation tasks. We evaluate state-of-the-art closed-source models, including GPT-4V, GPT-4o, Gemini-1.5-Pro,…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Yu Li , Yuchen Zheng , Giles Hamilton-Fletcher , Marco Mezzavilla , Yao Wang , Sundeep Rangan , Maurizio Porfiri , Zhou Yu , John-Ross Rizzo

Immersive technologies expand the potential for collaborative sense-making and visual analysis via head-worn displays (HWDs), offering customizable, high-resolution perspectives of a shared visualization space. In such an immersive…

人机交互 · 计算机科学 2025-11-25 Tamzid Hossain , Md. Fahimul Islam , Farida Chowdhury

Charts are ubiquitous as they help people understand and reason with data. Recently, various downstream tasks, such as chart question answering, chart2text, and fact-checking, have emerged. Large Vision-Language Models (LVLMs) show promise…

While shared autonomy offers significant potential for assistive robotics, key questions remain about how to effectively map 2D control inputs to 6D robot motions. An intuitive framework should allow users to input commands effortlessly,…

机器人学 · 计算机科学 2025-01-29 Shalutha Rajapakshe , Jean-Marc Odobez , Emmanuel Senft

This paper presents a novel framework for accessible and pedagogically-grounded robot explainability, designed to support human-robot interaction (HRI) with users who have diverse cognitive, communicative, or learning needs. We combine…

Online interactions and e-commerce are commonplace among BLV users. Despite the implementation of web accessibility standards, many e-commerce platforms continue to present challenges to screen reader users, particularly in areas like…

人机交互 · 计算机科学 2025-04-23 Yaman Yu , Bektur Ryskeldiev , Ayaka Tsutsui , Matthew Gillingham , Yang Wang

It is a challenging task for visually impaired people to perceive their surrounding environment due to the complexity of the natural scenes. Their personal and social activities are thus highly limited. This paper introduces a Large…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Zezhou Chen , Zhaoxiang Liu , Kai Wang , Kohou Wang , Shiguo Lian

Blind people have limited opportunities to explore an environment based on their interests. While existing navigation systems could provide them with surrounding information while navigating, they have limited scalability as they require…

人机交互 · 计算机科学 2025-02-14 Masaki Kuribayashi , Kohei Uehara , Allan Wang , Shigeo Morishima , Chieko Asakawa

Iteration of training and evaluating a machine learning model is an important process to improve its performance. However, while teachable interfaces enable blind users to train and test an object recognizer with photos taken in their…

The purpose of this study is to provide an accessibility measure of web-pages, in order to draw disabled users to the pages that have been designed to be ac-cessible to them. Our approach is based on the theory of belief functions, using…

人机交互 · 计算机科学 2015-01-21 Jean-Christophe Dubois , Yolande Le Gall , Arnaud Martin

Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks. Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs. To study the…

计算与语言 · 计算机科学 2025-06-09 Yingjie Zhu , Xuefeng Bai , Kehai Chen , Yang Xiang , Jun Yu , Min Zhang

Web accessibility aims to ensure that web content and services are usable by people with diverse abilities. In recent years, Large Language Models (LLMs) have been increasingly explored to support accessibility-related tasks on the web,…

数字图书馆 · 计算机科学 2026-05-15 Wajdi Aljedaani , Rubel Hassan Mollik

For people affected by blindness and low vision (BLV), safe and independent navigation remains a major challenge, impacting over 2.2 billion individuals worldwide. Although multimodal large language models (MLLMs) offer new opportunities…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Junhyeok Kim , Jaewoo Park , Junhee Park , Sangeyl Lee , Jiwan Chung , Jisung Kim , Ji Hoon Joung , Youngjae Yu

Diagrams represent a form of visual language that encodes abstract concepts and relationships through structured symbols and their spatial arrangements. Unlike natural images, they are inherently symbolic, and entirely artificial. They thus…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Yanpeng Sun , Shan Zhang , Wei Tang , Aotian Chen , Piotr Koniusz , Kai Zou , Yuan Xue , Anton van den Hengel

Digital geographic maps remain largely inaccessible to blind and low-vision individuals (BLVIs), despite global legislation adopting the Web Content Accessibility Guidelines (WCAG). A critical gap exists in defining "equivalent purpose" for…

人机交互 · 计算机科学 2026-05-15 Brandon Biggs , David Sloan , Brett Oppegaard , Nicholas A. Giudice , James M. Coughlan , Bruce N. Walker

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clarified to what extent LVLMs possess the ability to understand…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Kazuki Hayashi , Yusuke Sakai , Hidetaka Kamigaito , Katsuhiko Hayashi , Taro Watanabe

It is challenging for humans -- particularly those living with physical disabilities -- to control high-dimensional, dexterous robots. Prior work explores learning embedding functions that map a human's low-dimensional inputs (e.g., via a…

机器人学 · 计算机科学 2021-05-04 Siddharth Karamcheti , Albert J. Zhai , Dylan P. Losey , Dorsa Sadigh

When collaborating face-to-face, people commonly use the surfaces and spaces around them to perform sensemaking tasks, such as spatially organising documents, notes or images. However, when people collaborate remotely using desktop…

人机交互 · 计算机科学 2022-10-17 Ying Yang , Tim Dwyer , Michael Wybrow , Benjamin Lee , Maxime Cordeil , Mark Billinghurst , Bruce H. Thomas

Tactile graphics are essential for providing access to visual information for the 43 million people globally living with vision loss. Traditional methods for creating these graphics are labor-intensive and cannot meet growing demand. We…

计算机视觉与模式识别 · 计算机科学 2025-05-16 Adnan Khan , Alireza Choubineh , Mai A. Shaaban , Abbas Akkasi , Majid Komeili
‹ 上一页 1 8 9 10 下一页 ›