中文
相关论文

相关论文: Audio Description Customization

200 篇论文

Audio Descriptions (ADs) convey essential on-screen information, allowing visually impaired audiences to follow videos. To be effective, ADs must form a coherent sequence that helps listeners to visualise the unfolding scene, rather than…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Eshika Khandelwal , Junyu Xie , Tengda Han , Max Bain , Arsha Nagrani , Andrew Zisserman , Gül Varol , Makarand Tapaswi

Fully automated vehicles (FAVs) hold promise for enhancing the mobility of blind and low-vision (BLV) individuals. To understand the situated interaction needs of BLV passengers, we conducted six on-road, and in-lab focus groups with 16…

人机交互 · 计算机科学 2025-10-31 Zhengtao Ma , Rafael Gomez , Togtokhtur Batbold , Zishuo Zhu , Yueteng Yu , Ronald Schroeter

Audio Description (AD) plays a pivotal role as an application system aimed at guaranteeing accessibility in multimedia content, which provides additional narrations at suitable intervals to describe visual elements, catering specifically to…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Seon-Ho Lee , Jue Wang , David Fan , Zhikang Zhang , Linda Liu , Xiang Hao , Vimal Bhat , Xinyu Li

Introduction: Virtual audiovisual technology and its methodology has yet to be established for psychoacoustic research. This study examined the effects of different audiovisual conditions on preference when listening to multi-talker…

人机交互 · 计算机科学 2023-01-18 Gerard Llorach , Maartje M. E. Hendrikse , Giso Grimm , Volker Hohmann

Often, the needs and visual abilities differ between the annotator group and the end user group. Generating detailed diagram descriptions for blind and low-vision (BLV) users is one such challenging domain. Sighted annotators could describe…

人工智能 · 计算机科学 2025-03-18 Wan Ju Kang , Eunki Kim , Na Min An , Sangryul Kim , Haemin Choi , Ki Hoon Kwak , James Thorne

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of convolutional…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Sudha Krishnamurthy

Autism Spectrum Disorder (ASD) is neurodevelopmental condition characterized by social interaction and communication difficulties, along with narrow and repetitive interests. Being an spectrum disorder, ASD affects individuals with a large…

软件工程 · 计算机科学 2018-05-10 Roberto E. Lopez-Herrejon , Gerardo Herrera , Javier Sevilla

Blind and Low Vision (BLV) people have adopted AI-powered visual interpretation applications to address their daily needs. While these applications have been helpful, prior work has found that users remain unsatisfied by their frequent…

人机交互 · 计算机科学 2025-03-11 Ricardo E. Gonzalez Penuela , Ruiying Hu , Sharon Lin , Tanisha Shende , Shiri Azenkot

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Yuxin Guo , Shuailei Ma , Shijie Ma , Xiaoyi Bao , Chen-Wei Xie , Kecheng Zheng , Tingyu Weng , Siyang Sun , Yun Zheng , Wei Zou

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current…

多媒体 · 计算机科学 2025-05-29 Yong Ren , Chenxing Li , Le Xu , Hao Gu , Duzhen Zhang , Yujie Chen , Manjie Xu , Ruibo Fu , Shan Yang , Dong Yu

Despite advances in assistive technologies, Blind and Low-Vision (BLV) individuals continue to face challenges in understanding their surroundings. Delivering concise, useful, and timely scene descriptions for ambient perception remains a…

分布式、并行与集群计算 · 计算机科学 2026-03-17 Jacob Bradshaw , Mohsen Riahi Alam , Bhanuja Ainary , Minseo Kim , Mohsen Amini Salehi

Multimodal large language models (MLLMs) have been integrated into visual interpretation applications to support Blind and Low Vision (BLV) users because of their accuracy and ability to provide rich, human-like interpretations. However,…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Ricardo Gonzalez Penuela , Felipe Arias-Russi , Victor Capriles

Recent advances in Visual Anomaly Detection (VAD) have introduced sophisticated algorithms leveraging embeddings generated by pre-trained feature extractors. Inspired by these developments, we investigate the adaptation of such algorithms…

Understanding and managing data privacy in the digital world can be challenging for sighted users, let alone blind and low-vision (BLV) users. There is limited research on how BLV users, who have special accessibility needs, navigate data…

人机交互 · 计算机科学 2023-10-16 Yuanyuan Feng , Abhilasha Ravichander , Yaxing Yao , Shikun Zhang , Rex Chen , Shomir Wilson , Norman Sadeh

Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily lives. However, these models often produce errors that are difficult to detect without…

人机交互 · 计算机科学 2025-07-22 Meng Chen , Akhil Iyer , Amy Pavel

Online shopping has become a valuable modern convenience, but blind or low vision (BLV) users still face significant challenges using it, because of: 1) inadequate image descriptions and 2) the inability to filter large amounts of…

人机交互 · 计算机科学 2021-02-02 Ruolin Wang , Zixuan Chen , Mingrui "Ray" Zhang , Zhaoheng Li , Zhixiu Liu , Zihan Dang , Chun Yu , Xiang "Anthony" Chen

The objective of this paper is an automatic Audio Description (AD) model that ingests movies and outputs AD in text form. Generating high-quality movie AD is challenging due to the dependency of the descriptions on context, and the limited…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Tengda Han , Max Bain , Arsha Nagrani , Gül Varol , Weidi Xie , Andrew Zisserman

Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterpart, motion…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Yixuan Ren , Yang Zhou , Jimei Yang , Jing Shi , Difan Liu , Feng Liu , Mingi Kwon , Abhinav Shrivastava

GPS and smartphones enable users to place location-based annotations, capturing rich environmental context. Previous research demonstrates that blind and low vision (BLV) people can use annotations to explore unfamiliar areas. However,…

Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly when visual information is unstable or occluded due to continuous camera movement.…