English
Related papers

Related papers: SPICA: Interactive Video Content Exploration throu…

200 papers

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader…

Sound · Computer Science 2025-03-25 Yong Ren , Chenxing Li , Manjie Xu , Wei Liang , Yu Gu , Rilin Chen , Dong Yu

The Audio Description (AD) task aims to generate descriptions of visual elements for visually impaired individuals to help them access long-form video content, like movies. With video feature, text, character bank and context information as…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Hanlin Wang , Zhan Tong , Kecheng Zheng , Yujun Shen , Limin Wang

People with special needs like blind and visually impaired (BVI) people can particularly benefit from using voice assistants providing spoken information input and output in everyday life. However, it is crucial to understand their needs…

Human-Computer Interaction · Computer Science 2022-03-14 Christina Oumard , Julian Kreimeier , Timo Götzelmann

Visually Impaired Assistance (VIA) aims to automatically help the visually impaired (VI) handle daily activities. The advancement of VIA primarily depends on developments in Computer Vision (CV) and Natural Language Processing (NLP), both…

Computation and Language · Computer Science 2024-02-13 Yi Zhao , Yilin Zhang , Rong Xiang , Jing Li , Hillming Li

We introduce Semantic Parsing in Contextual Environments (SPICE), a task designed to enhance artificial agents' contextual awareness by integrating multimodal inputs with prior contexts. SPICE goes beyond traditional semantic parsing by…

Computation and Language · Computer Science 2024-06-11 Jordan Voas , Raymond Mooney , David Harwath

Visual impairment affects the ability of people to live a life like normal people. Such people face challenges in performing activities of daily living, such as reading, writing, traveling and participating in social gatherings. Many…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Mirza Samad Ahmed Baig , Syeda Anshrah Gillani , Shahid Munir Shah , Mahmoud Aljawarneh , Abdul Akbar Khan , Muhammad Hamzah Siddiqui

People who are blind or have low vision (BLV) encounter numerous challenges in their daily lives and work. To support them, various haptic assistive tools have been developed. Despite these advancements, the effective utilization of these…

Human-Computer Interaction · Computer Science 2024-12-30 Chutian Jiang , Emily Kuang , Mingming Fan

Prior research has shown how 'content preview tools' improve speed and accuracy of user relevance judgements across different information retrieval tasks. This paper describes a novel user interface tool, the Content Flow Bar, designed to…

Vision-language-action models (VLAs) have become an increasingly popular approach for addressing robot manipulation problems in recent years. However, such models need to output actions at a rate suitable for robot control, which limits the…

Robotics · Computer Science 2025-09-30 Eric Hannus , Miika Malin , Tran Nguyen Le , Ville Kyrki

Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capability that remains underexplored in existing vision-language systems. While prior work on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Jiangtao Wu , Shihao Li , Zhaozhou Bian , Jialu Chen , Runzhe Wen , An Ping , Yiwen He , Jiakai Wang , Yuanxing Zhang , Jiaheng Liu

This paper focuses on the challenge of answering questions in scenarios that are composed of rich and complex dynamic audio-visual components. Although existing Multimodal Large Language Models (MLLMs) can respond to audio-visual content,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Qilang Ye , Zitong Yu , Rui Shao , Xinyu Xie , Philip Torr , Xiaochun Cao

Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models, which, despite strong generalization in static manipulation, struggle in dynamic scenarios requiring rapid perception, temporal anticipation,…

Robotics · Computer Science 2026-01-30 Haozhe Xie , Beichen Wen , Jiarui Zheng , Zhaoxi Chen , Fangzhou Hong , Haiwen Diao , Ziwei Liu

One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectively embed spatio-temporal properties of visual scenes at the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Nikita Araslanov , Martin Sundermeyer , Hidenobu Matsuki , David Joseph Tan , Federico Tombari

Learning from audio-visual data offers many possibilities to express correspondence between the audio and visual content, similar to the human perception that relates aural and visual information. In this work, we present a method for…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-23 Shanshan Wang , Archontis Politis , Annamaria Mesaros , Tuomas Virtanen

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the remote user experience of a video conference by automatically panning, tilting and zooming of a video conferencing camera: users subjectively rate…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-26 Ross Cutler , Ramin Mehran , Sam Johnson , Cha Zhang , Adam Kirk , Oliver Whyte , Adarsh Kowdle

Most existing assistive navigation tools focus on providing real-time guidance for Blind and Low-Vision (BLV) people, but few support building a holistic spatial understanding of unfamiliar environments before travel. Such cognitive map…

Human-Computer Interaction · Computer Science 2026-04-17 Li Liu , Jiaming Qu , Marc Jowell Bagaoisan , David T. Lee , Leilani H. Gilpin

Video Question Answering (VidQA) exhibits remarkable potential in facilitating advanced machine reasoning capabilities within the domains of Intelligent Traffic Monitoring and Intelligent Transportation Systems. Nevertheless, the…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Ehsan Qasemi , Jonathan M. Francis , Alessandro Oltramari

Audio Description is a narrated commentary designed to aid vision-impaired audiences in perceiving key visual elements in a video. While short-form video understanding has advanced rapidly, a solution for maintaining coherent long-term…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Adrienne Deganutti , Simon Hadfield , Andrew Gilbert

Prompting and steering techniques are well established in general-purpose generative AI, yet assistive visual question answering (VQA) tools for blind users still follow rigid interaction patterns with limited opportunities for…

Human-Computer Interaction · Computer Science 2026-02-20 Farnaz Zamiri Zeraati , Yang Trista Cao , Yuehan Qiao , Hal Daumé , Hernisa Kacorri

Real-world videos often show routine activities punctuated by memorable, surprising events. However, most Video-LLMs process videos by sampling frames uniformly, likely missing critical moments that define a video's narrative. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Sahithya Ravi , Aditya Chinchure , Raymond T. Ng , Leonid Sigal , Vered Shwartz
‹ Prev 1 4 5 6 7 8 10 Next ›