中文
相关论文

相关论文: Audio Description Customization

200 篇论文

Robotic guide dogs hold significant potential to enhance the autonomy and mobility of blind or visually impaired (BVI) individuals by offering universal assistance over unstructured terrains at affordable costs. However, the design of…

机器人学 · 计算机科学 2025-01-09 J. Taery Kim , Morgan Byrd , Jack L. Crandell , Bruce N. Walker , Greg Turk , Sehoon Ha

Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information. However, a thorough understanding of videos significantly…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yudong Yang , Jimin Zhuang , Guangzhi Sun , Changli Tang , Yixuan Li , Peihan Li , Yifan Jiang , Wei Li , Zejun Ma , Chao Zhang

Vision-Language Models (VLMs) are increasingly used by blind and low-vision (BLV) people to identify and understand products in their everyday lives, such as food, personal care items, and household goods. Despite their prevalence, we lack…

人机交互 · 计算机科学 2026-04-01 Kapil Garg , Xinru Tang , Jimin Heo , Dwayne R. Morgan , Darren Gergle , Erik B. Sudderth , Anne Marie Piper

We present AVID, the first large-scale benchmark for audio-visual inconsistency understanding in videos. While omni-modal large language models excel at temporally aligned tasks such as captioning and question answering, they struggle to…

多媒体 · 计算机科学 2026-04-16 Zixuan Chen , Depeng Wang , Hao Lin , Li Luo , Ke Xu , Ya Guo , Huijia Zhu , Tanfeng Sun , Xinghao Jiang

Content-aware streaming requires dynamic, chunk-level importance weights to optimize subjective quality of experience (QoE). However, direct human annotation is prohibitively expensive while vision-saliency models generalize poorly. We…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Jiahui Chen , Bo Peng , Lianchen Jia , Zeyu Zhang , Tianchi Huang , Lifeng Sun

Visually impaired people face significant challenges when attempting to interact with and understand complex environments, and traditional assistive technologies often struggle to quickly provide necessary contextual understanding and…

人机交互 · 计算机科学 2025-05-02 Bhanuja Ainary

People with special needs like blind and visually impaired (BVI) people can particularly benefit from using voice assistants providing spoken information input and output in everyday life. However, it is crucial to understand their needs…

人机交互 · 计算机科学 2022-03-14 Christina Oumard , Julian Kreimeier , Timo Götzelmann

Automated live visual descriptions can aid blind people in understanding their surroundings with autonomy and independence. However, providing descriptions that are rich, contextual, and just-in-time has been a long-standing challenge in…

人机交互 · 计算机科学 2024-08-14 Ruei-Che Chang , Yuxuan Liu , Anhong Guo

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Ruohan Gao , Kristen Grauman

Abbreviation expansion is a strategy used to speed up communication by limiting the amount of typing and using a language model to suggest expansions. Here we look at personalizing a Large Language Model's (LLM) suggestions based on prior…

计算与语言 · 计算机科学 2023-12-25 Katrin Tomanek , Shanqing Cai , Subhashini Venugopalan

The exponential growth of short-video content has ignited a surge in the necessity for efficient, automated solutions to video editing, with challenges arising from the need to understand videos and tailor the editing according to user…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Dabing Cheng , Haosen Zhan , Xingchen Zhao , Guisheng Liu , Zemin Li , Jinghui Xie , Zhao Song , Weiguo Feng , Bingyue Peng

Audio captioning aims to generate text descriptions of audio clips. In the real world, many objects produce similar sounds. How to accurately recognize ambiguous sounds is a major challenge for audio captioning. In this work, inspired by…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Xubo Liu , Qiushi Huang , Xinhao Mei , Haohe Liu , Qiuqiang Kong , Jianyuan Sun , Shengchen Li , Tom Ko , Yu Zhang , Lilian H. Tang , Mark D. Plumbley , Volkan Kılıç , Wenwu Wang

Recent advancements in Vision-Language (VL) models have sparked interest in their deployment on edge devices, yet challenges in handling diverse visual modalities, manual annotation, and computational constraints remain. We introduce…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Kaiwen Cai , Zhekai Duan , Gaowen Liu , Charles Fleming , Chris Xiaoxuan Lu

Augmented, virtual and mixed reality technologies offer new ways of interacting with digital media. However, such technologies are not well explored for people with different ranges of abilities beyond a few specific navigation and gaming…

人机交互 · 计算机科学 2021-01-11 Pradipta Biswas , Pilar Orero , Manohar Swaminathan , Kavita Krishnaswamy , Peter Robinson

Large Language Models (LLMs) have shown remarkable potential in recommending everyday actions as personal AI assistants, while Explainable AI (XAI) techniques are being increasingly utilized to help users understand why a recommendation is…

Blind and low-vision (BLV) users remain largely excluded from three-dimensional (3D) surface and point data visualizations due to the reliance on visual interaction. Existing approaches inadequately support non-visual access, especially in…

人机交互 · 计算机科学 2025-08-13 Sanchita S. Kamath , Aziz N. Zeidieh , JooYoung Seo

Audio-visual speech recognition (AVSR) aims to transcribe human speech using both audio and video modalities. In practical environments with noise-corrupted audio, the role of video information becomes crucial. However, prior works have…

音频与语音处理 · 电气工程与系统科学 2024-10-15 Sungnyun Kim , Kangwook Jang , Sangmin Bae , Hoirin Kim , Se-Young Yun

The overview-detail design pattern, characterized by an overview of multiple items and a detailed view of a selected item, is ubiquitously implemented across software interfaces. Designers often try to account for all users, but ultimately…

人机交互 · 计算机科学 2025-03-12 Bryan Min , Allen Chen , Yining Cao , Haijun Xia

Visual dubbing, the synchronization of facial movements with new speech, is crucial for making content accessible across different languages, enabling broader global reach. However, current methods face significant limitations. Existing…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Binyamin Manela , Sharon Gannot , Ethan Fetyaya

Automated driving system - dedicated vehicles (ADS-DVs), specially designed for people with various disabilities, can be beneficial to improve their mobility. However, research related to autonomous vehicles (AVs) for people with cognitive…