English
Related papers

Related papers: When Vision Speaks for Sound

200 papers

While natural-language explanations from large language models (LLMs) are widely adopted to improve transparency and trust, their impact on objective human-AI team performance remains poorly understood. We identify a Persuasion Paradox:…

Human-Computer Interaction · Computer Science 2026-04-07 Ruth Cohen , Lu Feng , Ayala Bloch , Sarit Kraus

Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic information remains limited. To systematically quantify this issue, we introduce…

Sound · Computer Science 2026-05-28 Jiacheng Pang , Ashutosh Chaubey , Mohammad Soleymani

In audio-visual navigation (AVN), an intelligent agent needs to navigate to a constantly sound-making object in complex 3D environments based on its audio and visual perceptions. While existing methods attempt to improve the navigation…

Sound · Computer Science 2022-06-02 Shunqi Mao , Chaoyi Zhang , Heng Wang , Weidong Cai

In this paper, we propose to make a systematic study on machines multisensory perception under attacks. We use the audio-visual event recognition task against multimodal adversarial attacks as a proxy to investigate the robustness of…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Yapeng Tian , Chenliang Xu

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far…

Sound · Computer Science 2025-03-31 Yunming Liang , Zihao Chen , Chaofan Ding , Xinhan Di

In this paper, we explore self-supervised audio-visual models that learn from instructional videos. Prior work has shown that these models can relate spoken words and sounds to visual content after training on a large-scale dataset of…

Digital video is central to communication, education, and entertainment, but without audio description (AD), blind and low-vision users are excluded. While crowdsourced platforms and vision-language models (VLMs) expand AD production,…

Human-Computer Interaction · Computer Science 2026-05-08 Lana Do , Gio Jung , Juvenal Francisco Barajas , Andrew Taylor Scott , Shasta Ihorn , Alexander Mario Blum , Vassilis Athitsos , Ilmi Yoon

The hallucination of large multimodal models (LMMs), providing responses that appear correct but are actually incorrect, limits their reliability and applicability. This paper aims to study the hallucination problem of LMMs in video…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Hongcheng Gao , Jiashu Qu , Jingyi Tang , Baolong Bi , Yue Liu , Hongyu Chen , Li Liang , Li Su , Qingming Huang

In the face of the video data deluge, today's expensive clip-level classifiers are increasingly impractical. We propose a framework for efficient action recognition in untrimmed video that uses audio as a preview mechanism to eliminate both…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Ruohan Gao , Tae-Hyun Oh , Kristen Grauman , Lorenzo Torresani

Deep learning has made significant impacts on multi-view stereo systems. State-of-the-art approaches typically involve building a cost volume, followed by multiple 3D convolution operations to recover the input image's pixel-wise depth.…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Zhenpei Yang , Zhile Ren , Qi Shan , Qixing Huang

Automatic transcriptions of consumer-generated multi-media content such as "Youtube" videos still exhibit high word error rates. Such data typically occupies a very broad domain, has been recorded in challenging conditions, with cheap…

Computation and Language · Computer Science 2017-12-08 Abhinav Gupta , Yajie Miao , Leonardo Neves , Florian Metze

Auditory foundation models, including auditory large language models (LLMs), process all sound inputs equally, independent of listener perception. However, human auditory perception is inherently selective: listeners focus on specific…

Augmented reality (AR) is emerging in visual search tasks for increasingly immersive interactions with virtual objects. We propose an AR approach providing visual and audio hints along with gaze-assisted instant post-task feedback for…

Human-Computer Interaction · Computer Science 2023-11-15 Yuchong Zhang , Adam Nowak , Yueming Xuan , Andrzej Romanowski , Morten Fjeld

Despite recent advances in text-to-speech (TTS) models, audio-visual-to-audio-visual (AV2AV) translation still faces a critical challenge: maintaining speaker consistency between the original and translated vocal and facial features. To…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-31 Sungwoo Cho , Jeongsoo Choi , Sungnyun Kim , Se-Young Yun

Vision-Language Models (VLMs) have been shown to be blind, often underutilizing their visual inputs even on tasks that require visual reasoning. In this work, we demonstrate that VLMs are selectively blind. They modulate the amount of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Wan-Cyuan Fan , Jiayun Luo , Declan Kutscher , Leonid Sigal , Ritwik Gupta

Multi-modal Large Language Models (MLLMs) have significantly advanced video reasoning, yet Video Question Answering (VideoQA) remains challenging due to its demand for temporal causal reasoning and evidence-grounded answer generation.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Kaixin zhang , Xiaohe Li , Jiahao Li , Haohua Wu , Xinyu Zhao , Zide Fan , Lei Wang

Recent Multimodal Large Language Models (MLLMs) excel on benchmark vision-language tasks, yet little is known about how input visual quality shapes their responses. Does higher perceptual quality of images already translate to better MLLM…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Shuo Xing , Lanqing Guo , Hongyuan Hua , Seoyoung Lee , Peiran Li , Yufei Wang , Zhangyang Wang , Zhengzhong Tu

Answering questions about images often requires combining visual understanding with external knowledge. Multimodal Large Language Models (MLLMs) provide a natural framework for this setting, but they often struggle to identify the most…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Marco Morini , Sara Sarto , Marcella Cornia , Lorenzo Baraldi

The goal of this paper is to learn strong lip reading models that can recognise speech in silent videos. Most prior works deal with the open-set visual speech recognition problem by adapting existing automatic speech recognition techniques…

Computer Vision and Pattern Recognition · Computer Science 2021-12-06 K R Prajwal , Triantafyllos Afouras , Andrew Zisserman

Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio…

Sound · Computer Science 2021-05-14 Xinyuan Qian , Maulik Madhavi , Zexu Pan , Jiadong Wang , Haizhou Li
‹ Prev 1 4 5 6 7 8 10 Next ›