English
Related papers

Related papers: SPICA: Interactive Video Content Exploration throu…

200 papers

Long video understanding is a significant and ongoing challenge in the intersection of multimedia and artificial intelligence. Employing large language models (LLMs) for comprehending video becomes an emerging and promising method. However,…

Computation and Language · Computer Science 2024-08-27 Yunxin Li , Xinyu Chen , Baotain Hu , Min Zhang

Video-grounded dialogues are very challenging due to (i) the complexity of videos which contain both spatial and temporal variations, and (ii) the complexity of user utterances which query different segments and/or different objects in…

Computer Vision and Pattern Recognition · Computer Science 2020-10-21 Hung Le , Doyen Sahoo , Nancy F. Chen , Steven C. H. Hoi

State-of-the-art Active Speaker Detection (ASD) approaches heavily rely on audio and facial features to perform, which is not a sustainable approach in wild scenarios. Although these methods achieve good results in the standard…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Tiago Roxo , Joana C. Costa , Pedro R. M. Inácio , Hugo Proença

Many blind and low vision (BLV) people are excluded from professional roles that may involve visual tasks due to access barriers and persisting stigmas. Advancing generative AI systems can support BLV people through providing contextual and…

Human-Computer Interaction · Computer Science 2025-10-13 Lucy Jiang , Lotus Zhang , Leah Findlater

The visual quality of an image is confounded by a number of intertwined factors including its semantic content, distortion characteristics and appearance properties such as brightness, contrast, sharpness, and colourfulness. Distilling high…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Fei Zhou , Tianhao Gu , Zhicong Huang , Guoping Qiu

Blind and low vision (BLV) internet users access images on the web via text descriptions. New vision-to-language models such as GPT-V, Gemini, and LLaVa can now provide detailed image descriptions on-demand. While prior research and…

Human-Computer Interaction · Computer Science 2024-09-06 Ananya Gubbi Mohanbabu , Amy Pavel

Audio description (AD) makes video content accessible to blind and low-vision (BLV) audiences, but producing high-quality descriptions is resource-intensive. Automated AD offers scalability, and prior studies show human-in-the-loop editing…

Human-Computer Interaction · Computer Science 2026-02-04 Lana Do , Shasta Ihorn , Charity Pitcher-Cooper , Juvenal Francisco Barajas , Gio Jung , Xuan Duy Anh Nguyen , Sanjay Mirani , Ilmi Yoon

Video accessibility is crucial for blind and low vision users for equitable engagements in education, employment, and entertainment. Despite the availability of professional and amateur services and tools, most human-generated descriptions…

Human-Computer Interaction · Computer Science 2022-01-12 Shasta Ihorn , Yue-Ting Siu , Aditya Bodi , Lothar Narins , Jose M. Castanon , Yash Kant , Abhishek Das , Ilmi Yoon , Pooyan Fazli

Blind and low vision (BLV) users often rely on alt text to understand what a digital image is showing. However, recent research has investigated how touch-based image exploration on touchscreens can supplement alt text. Touchscreen-based…

Human-Computer Interaction · Computer Science 2023-02-21 Vishnu Nair , Hanxiu 'Hazel' Zhu , Brian A. Smith

Video-based spatial cognition is vital for robotics and embodied AI but challenges current Vision-Language Models (VLMs). This paper makes two key contributions. First, we introduce ViCA (Visuospatial Cognitive Assistant)-322K, a diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Qi Feng

Audio description, a form of trans-modal media translation, allows people who are blind or visually impaired access to visually-oriented, socio-cultural, or historical public discourse alike. Although audio description has gained more…

Human-Computer Interaction · Computer Science 2018-09-18 Philipp Jordan , Brett Oppegaard

Audio descriptions make videos accessible to those who cannot see them by describing visual content in audio. Producing audio descriptions is challenging due to the synchronous nature of the audio description that must fit into gaps of…

Human-Computer Interaction · Computer Science 2020-10-09 Amy Pavel , Gabriel Reyes , Jeffrey P. Bigham

This paper presents AIDEN, an artificial intelligence-based assistant designed to enhance the autonomy and daily quality of life of visually impaired individuals, who often struggle with object identification, text reading, and navigation…

Spatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized equipment and skills, posing a high barrier for amateur video creators. We…

Human-Computer Interaction · Computer Science 2024-04-24 Zheng Ning , Zheng Zhang , Jerrick Ban , Kaiwen Jiang , Ruohong Gan , Yapeng Tian , Toby Jia-Jun Li

Referring Image Segmentation (RIS) aims to segment a target object described by a natural language expression. Existing methods have evolved by leveraging the vision information into the language tokens. To more effectively exploit visual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Yubin Cho , Hyunwoo Yu , Kyeongbo Kong , Kyomin Sohn , Bongjoon Hyun , Suk-Ju Kang

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 Rahul Sharma , Krishna Somandepalli , Shrikanth Narayanan

Videos offer rich audiovisual information that can support people in performing activities of daily living (ADLs), but they remain largely inaccessible to blind or low-vision (BLV) individuals. In cooking, BLV people often rely on…

Human-Computer Interaction · Computer Science 2025-07-16 Zheng Ning , Leyang Li , Daniel Killough , JooYoung Seo , Patrick Carrington , Yapeng Tian , Yuhang Zhao , Franklin Mingzhe Li , Toby Jia-Jun Li

In this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically strategize the utilization of external tools and to investigate their…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Ziniu Hu , Ahmet Iscen , Chen Sun , Kai-Wei Chang , Yizhou Sun , David A Ross , Cordelia Schmid , Alireza Fathi

In this paper we explore the opportunities brought by cognitive augmentation to provide a more natural and accessible web browsing experience. We explore these opportunities through \textit{conversational web browsing}, an emerging…

Computers and Society · Computer Science 2020-12-08 Alessandro Pina , Marcos Baez , Florian Daniel

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Chaoyu Li , Sid Padmanabhuni , Maryam Cheema , Hasti Seifi , Pooyan Fazli