English
Related papers

Related papers: VGGSounder: Audio-Visual Evaluations for Foundatio…

200 papers

The success of vision-language models is primarily attributed to effective alignment across modalities such as vision and language. However, modality gaps persist in existing alignment algorithms and appear necessary for human perception as…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Hanqi Yan , Xiangxiang Cui , Lu Yin , Jindong Gu , Paul Pu Liang , Yulan He , Yifei Wang

As audio/visual classification models are widely deployed for sensitive tasks like content filtering at scale, it is critical to understand their robustness along with improving the accuracy. This work aims to study several key questions…

Computer Vision and Pattern Recognition · Computer Science 2020-11-17 Juncheng B Li , Kaixin Ma , Shuhui Qu , Po-Yao Huang , Florian Metze

In this paper, we propose a multi-level attention model to solve the weakly labelled audio classification problem. The objective of audio classification is to predict the presence or absence of audio events in an audio clip. Recently,…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-08 Changsong Yu , Karim Said Barsim , Qiuqiang Kong , Bin Yang

The rapid development of generative audio raises ethical and security concerns stemming from forged data, making deepfake sound detection an important safeguard against the malicious use of such technologies. Although prior studies have…

Sound · Computer Science 2025-09-29 Zeyu Xie , Yaoyun Zhang , Xuenan Xu , Yongkang Yin , Chenxing Li , Mengyue Wu , Yuexian Zou

Audio-Visual Semantic Segmentation (AVSS) aligns audio and video at the pixel level but requires costly per-frame annotations. We introduce Weakly Supervised Audio-Visual Semantic Segmentation (WSAVSS), which uses only video-level labels to…

Multimedia · Computer Science 2026-03-24 Chengzhi Li , Heyan Huang , Ping Jian , Yanghao Zhou

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak

The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years. However, researchers face challenges due to the costly and time-consuming collection process of existing audio-language datasets, which…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-22 Xinhao Mei , Chutong Meng , Haohe Liu , Qiuqiang Kong , Tom Ko , Chengqi Zhao , Mark D. Plumbley , Yuexian Zou , Wenwu Wang

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

Sound source localization is a typical and challenging task that predicts the location of sound sources in a video. Previous single-source methods mainly used the audio-visual association as clues to localize sounding objects in each image.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Shentong Mo , Yapeng Tian

Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Yilin Ye , Shishi Xiao , Xingchen Zeng , Wei Zeng

Learning from multiple modalities, such as audio and video, offers opportunities for leveraging complementary information, enhancing robustness, and improving contextual understanding and performance. However, combining such modalities…

Multimedia · Computer Science 2024-10-15 Konstantinos Kontras , Christos Chatzichristos , Matthew Blaschko , Maarten De Vos

Human perception and experience of music is highly context-dependent. Contextual variability contributes to differences in how we interpret and interact with music, challenging the design of robust models for information retrieval.…

Sound · Computer Science 2022-10-31 Kleanthis Avramidis , Shanti Stewart , Shrikanth Narayanan

The performance of the keyword spotting (KWS) system based on audio modality, commonly measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. Therefore, audio-visual keyword spotting,…

Sound · Computer Science 2023-03-15 Ao Zhang , He Wang , Pengcheng Guo , Yihui Fu , Lei Xie , Yingying Gao , Shilei Zhang , Junlan Feng

Weakly supervised visual grounding (VG) aims to locate objects in images based on text descriptions. Despite significant progress, existing methods lack strong cross-modal reasoning to distinguish subtle semantic differences in text…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yidan Wang , Chenyi Zhuang , Wutao Liu , Pan Gao , Nicu Sebe

Audio-visual learning helps to comprehensively understand the world by fusing practical information from multiple modalities. However, recent studies show that the imbalanced optimization of uni-modal encoders in a joint-learning model is a…

Sound · Computer Science 2023-03-14 Ruize Xu , Ruoxuan Feng , Shi-Xiong Zhang , Di Hu

Audio tagging is the task of predicting the presence or absence of sound classes within an audio clip. Previous work in audio tagging focused on relatively small datasets limited to recognising a small number of sound classes. We…

Sound · Computer Science 2019-12-11 Qiuqiang Kong , Changsong Yu , Turab Iqbal , Yong Xu , Wenwu Wang , Mark D. Plumbley

We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given two images and two captions, the goal is to match them…

Computer Vision and Pattern Recognition · Computer Science 2022-04-26 Tristan Thrush , Ryan Jiang , Max Bartolo , Amanpreet Singh , Adina Williams , Douwe Kiela , Candace Ross

We introduce VisoGender, a novel dataset for benchmarking gender bias in vision-language models. We focus on occupation-related biases within a hegemonic system of binary gender, inspired by Winograd and Winogender schemas, where each image…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Siobhan Mackenzie Hall , Fernanda Gonçalves Abrantes , Hanwen Zhu , Grace Sodunke , Aleksandar Shtedritski , Hannah Rose Kirk

The analysis, processing, and extraction of meaningful information from sounds all around us is the subject of the broader area of audio analytics. Audio captioning is a recent addition to the domain of audio analytics, a cross-modal…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-04 Sandeep Kothinti , Dimitra Emmanouilidou

The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal understanding, enabling more sophisticated and accurate integration of visual and textual information across various tasks, including image and video…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Hang Hua , Yunlong Tang , Ziyun Zeng , Liangliang Cao , Zhengyuan Yang , Hangfeng He , Chenliang Xu , Jiebo Luo