English
Related papers

Related papers: Mitigating Audiovisual Mismatch in Visual-Guide Au…

200 papers

We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal…

Multimedia · Computer Science 2025-07-28 Hyunwoo Oh , SeungJu Cha , Kwanyoung Lee , Si-Woo Kim , Dong-Jin Kim

Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complementary and…

Sound · Computer Science 2023-09-29 R. Gnana Praveen , Jahangir Alam

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of…

Computer Vision and Pattern Recognition · Computer Science 2021-12-15 Yidi Li , Hong Liu , Hao Tang

Image captioning aims at generating descriptive and meaningful textual descriptions of images, enabling a broad range of vision-language applications. Prior works have demonstrated that harnessing the power of Contrastive Image Language…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 Longtian Qiu , Shan Ning , Xuming He

Content-based music information retrieval has seen rapid progress with the adoption of deep learning. Current approaches to high-level music description typically make use of classification models, such as in auto-tagging or genre and mood…

Sound · Computer Science 2021-12-09 Ilaria Manco , Emmanouil Benetos , Elio Quinton , Gyorgy Fazekas

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

Vision-language models have been key to the development of open-vocabulary 2D semantic segmentation. Lifting these models from 2D images to 3D scenes, however, remains a challenging problem. Existing approaches typically back-project and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Tomas Berriel Martins , Martin R. Oswald , Javier Civera

Multimodal sentiment analysis (MSA) leverages information fusion from diverse modalities (e.g., text, audio, visual) to enhance sentiment prediction. However, simple fusion techniques often fail to account for variations in modality…

Machine Learning · Computer Science 2025-10-03 Han Wu , Yanming Sun , Yunhe Yang , Derek F. Wong

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Trevine Oorloff , Surya Koppisetti , Nicolò Bonettini , Divyaraj Solanki , Ben Colman , Yaser Yacoob , Ali Shahriyari , Gaurav Bharaj

Visual captioning aims to generate textual descriptions given images or videos. Traditionally, image captioning models are trained on human annotated datasets such as Flickr30k and MS-COCO, which are limited in size and diversity. This…

Computer Vision and Pattern Recognition · Computer Science 2021-03-01 Marimuthu Kalimuthu , Aditya Mogadala , Marius Mosbach , Dietrich Klakow

Machines that can represent and describe environmental soundscapes have practical potential, e.g., for audio tagging and captioning systems. Prevailing learning paradigms have been relying on parallel audio-text data, which is, however,…

Sound · Computer Science 2022-05-04 Yanpeng Zhao , Jack Hessel , Youngjae Yu , Ximing Lu , Rowan Zellers , Yejin Choi

Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions and contexts. This…

Computation and Language · Computer Science 2025-01-07 Ariel Shaulov , Tal Shaharabany , Eitan Shaar , Gal Chechik , Lior Wolf

Many previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse musical video data, voice activity detection is a necessary…

Sound · Computer Science 2021-06-23 Yuanbo Hou , Zhesong Yu , Xia Liang , Xingjian Du , Bilei Zhu , Zejun Ma , Dick Botteldooren

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges: (1) the scarcity…

Sound · Computer Science 2026-04-30 Yusheng Dai , Zehua Chen , Yuxuan Jiang , Baolong Gao , Qiuhong Ke , Jianfei Cai , Jun Zhu

Effectively aligning with human judgment when evaluating machine-generated image captions represents a complex yet intriguing challenge. Existing evaluation metrics like CIDEr or CLIP-Score fall short in this regard as they do not take into…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Sara Sarto , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Xinlong Chen , Yue Ding , Weihong Lin , Jingyun Hua , Linli Yao , Yang Shi , Bozhou Li , Yuanxing Zhang , Qiang Liu , Pengfei Wan , Liang Wang , Tieniu Tan

Multimodal models play a key role in empathy detection, but their performance can suffer when modalities provide conflicting cues. To understand these failures, we examine cases where unimodal and multimodal predictions diverge. Using…

Computation and Language · Computer Science 2025-11-12 Maya Srikanth , Run Chen , Julia Hirschberg

Text-image cross-modal retrieval is a challenging task in the field of language and vision. Most previous approaches independently embed images and sentences into a joint embedding space and compare their similarities. However, previous…

Computer Vision and Pattern Recognition · Computer Science 2019-09-13 Zihao Wang , Xihui Liu , Hongsheng Li , Lu Sheng , Junjie Yan , Xiaogang Wang , Jing Shao

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovisual fusion by…