English
Related papers

Related papers: DanmuA11y: Making Time-Synced On-Screen Video Comm…

200 papers

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Chaoyu Li , Sid Padmanabhuni , Maryam Cheema , Hasti Seifi , Pooyan Fazli

Video-sharing platforms (VSPs) have been increasingly embracing social features such as likes, comments, and Danmaku to boost user engagement. However, viewers may post inappropriate content through video commentary to gain attention or…

Human-Computer Interaction · Computer Science 2024-11-08 Siying Hu , Zhicong Lu

Danmaku, users' live comments synchronized with, and overlaying on videos, has recently shown potential in promoting online video-based learning. However, user-generated danmaku can be scarce-especially in newer or less viewed videos and…

Human-Computer Interaction · Computer Science 2025-04-28 Zipeng Ji , Pengcheng An , Jian Zhao

On general video-sharing platforms like YouTube, comments are displayed independently of video playback. As viewers often read comments while watching a video, they may encounter ones referring to moments unrelated to the current scene,…

Multimedia · Computer Science 2026-03-30 Minsun Kim , Dawon Lee , Junyong Noh

Previous research underscored the potential of danmaku--a text-based commenting feature on videos--in engaging hearing audiences. Yet, for many Deaf and hard-of-hearing (DHH) individuals, American Sign Language (ASL) takes precedence over…

Human-Computer Interaction · Computer Science 2024-03-28 Si Chen , Haocong Cheng , Jason Situ , Desirée Kirst , Suzy Su , Saumya Malhotra , Lawrence Angrave , Qi Wang , Yun Huang

While audio description (AD) is the standard approach for making videos accessible to blind and low vision (BLV) people, existing AD guidelines do not consider BLV users' varied preferences across viewing scenarios. These scenarios range…

Human-Computer Interaction · Computer Science 2024-03-19 Lucy Jiang , Crescentia Jung , Mahika Phutane , Abigale Stangl , Shiri Azenkot

Authors make their videos visually accessible by adding audio descriptions (AD), and auditorily accessible by adding closed captions (CC). However, creating AD and CC is challenging and tedious, especially for non-professional describers…

Human-Computer Interaction · Computer Science 2025-02-19 Xingyu "Bruce" Liu , Ruolin Wang , Dingzeyu Li , Xiang 'Anthony' Chen , Amy Pavel

Videos make exercise instruction widely available, but they rely on visual demonstrations that blind and low vision (BLV) learners cannot see. While audio descriptions (AD) can make videos accessible, describing movements remains…

Human-Computer Interaction · Computer Science 2026-03-02 Ujjaini Das , Shreya Kappala , Meng Chen , Mina Huh , Amy Pavel

Video content remains largely inaccessible to blind and low-vision (BLV) users. To address this, we introduce a prototype that leverages a multimodal agent - powered by a novel conversational architecture using a multimodal large language…

User emotion analysis toward videos is to automatically recognize the general emotional status of viewers from the multimedia content embedded in the online video stream. Existing works fall in two categories: 1) visual-based methods, which…

Computer Vision and Pattern Recognition · Computer Science 2018-11-20 Chenchen Li , Jialin Wang , Hongwei Wang , Miao Zhao , Wenjie Li , Xiaotie Deng

Online video platforms have gained increased popularity due to their ability to support information consumption and sharing and the diverse social interactions they afford. Danmaku, a real-time commentary feature that overlays user comments…

Human-Computer Interaction · Computer Science 2025-02-11 Siying Hu , Huanchen Wang , Yu Zhang , Piaohong Wang , Zhicong Lu

Presenters commonly use slides as visual aids for informative talks. When presenters fail to verbally describe the content on their slides, blind and visually impaired audience members lose access to necessary content, making the…

Human-Computer Interaction · Computer Science 2021-03-29 Yi-Hao Peng , JiWoong Jang , Jeffrey P. Bigham , Amy Pavel

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Yuxin Guo , Shuailei Ma , Shijie Ma , Xiaoyi Bao , Chen-Wei Xie , Kecheng Zheng , Tingyu Weng , Siyang Sun , Yun Zheng , Wei Zou

Effective visual accessibility in Virtual Reality (VR) is crucial for Blind and Low Vision (BLV) users. However, designing visual accessibility systems is challenging due to the complexity of 3D VR environments and the need for techniques…

Human-Computer Interaction · Computer Science 2025-02-07 Junlong Chen , Rosella P. Galindo Esparza , Vanja Garaj , Per Ola Kristensson , John Dudley

We present DyMU, an efficient, training-free framework that dynamically reduces the computational burden of vision-language models (VLMs) while maintaining high task performance. Our approach comprises two key components. First, Dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Zhenhailong Wang , Senthil Purushwalkam , Caiming Xiong , Silvio Savarese , Heng Ji , Ran Xu

As virtual 3D environments become more prevalent, equitable access is essential for blind and low-vision (BLV) users, who face challenges with spatial awareness, navigation, and interaction. Prior work has explored supplementing visual…

Human-Computer Interaction · Computer Science 2026-02-10 Xinyun Cao , Kexin Phyllis Ju , Chenglin Li , Venkatesh Potluri , Dhruv Jain

Audio Description (AD) provides essential access to visual media for blind and low vision (BLV) audiences. Yet current AD production tools remain largely inaccessible to BLV video creators, who possess valuable expertise but face barriers…

Human-Computer Interaction · Computer Science 2026-02-10 Franklin Mingzhe Li , Michael Xieyang Liu , Cynthia L. Bennett , Shaun K. Kane

Modern Vision-Language Models (VLMs) achieve impressive performance but are limited by the quadratic complexity of self-attention, which prevents their deployment on edge devices and makes their understanding of high-resolution images and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Hongjie Wang , Niraj K. Jha

Audio description (AD) makes video content accessible to blind and low-vision (BLV) audiences, but producing high-quality descriptions is resource-intensive. Automated AD offers scalability, and prior studies show human-in-the-loop editing…

Human-Computer Interaction · Computer Science 2026-02-04 Lana Do , Shasta Ihorn , Charity Pitcher-Cooper , Juvenal Francisco Barajas , Gio Jung , Xuan Duy Anh Nguyen , Sanjay Mirani , Ilmi Yoon

The rapid growth of online video content has outpaced efforts to make visual information accessible to blind and low vision (BLV) audiences. While professional Audio Description (AD) remains the gold standard, it is costly and difficult to…

Human-Computer Interaction · Computer Science 2025-08-13 Ruolin wang , Xingyu Liu , Biao Wang , Wayne Zhang , Ziqian Liao , Ziwen Li , Amy Pavel , Xiang 'Anthony' Chen
‹ Prev 1 2 3 10 Next ›