English
Related papers

Related papers: VGGSounder: Audio-Visual Evaluations for Foundatio…

200 papers

Advanced Audio-Visual Speech Recognition (AVSR) systems have been observed to be sensitive to missing video frames, performing even worse than single-modality models. While applying the dropout technique to the video modality enhances…

Sound · Computer Science 2024-03-08 Yusheng Dai , Hang Chen , Jun Du , Ruoyu Wang , Shihao Chen , Jiefeng Ma , Haotian Wang , Chin-Hui Lee

Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Hongyu Wang , Hui Li , Bo Li

Audio-language models have shown promising results in various sound understanding tasks, yet they remain limited in their ability to reason over the fine-grained semantics of sound. In this paper, we present AudSemThinker, a model whose…

Sound · Computer Science 2025-10-01 Gijs Wijngaard , Elia Formisano , Michele Esposito , Michel Dumontier

A versatile medical image segmentation model applicable to images acquired with diverse equipment and protocols can facilitate model deployment and maintenance. However, building such a model typically demands a large, diverse, and fully…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Xiaoyang Chen , Hao Zheng , Yuemeng Li , Yuncong Ma , Liang Ma , Hongming Li , Yong Fan

We propose an Explicit Conditional Multimodal Variational Auto-Encoder (ECMVAE) for audio-visual segmentation (AVS), aiming to segment sound sources in the video sequence. Existing AVS methods focus on implicit feature fusion strategies,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Yuxin Mao , Jing Zhang , Mochu Xiang , Yiran Zhong , Yuchao Dai

Recent advancements in multimodal techniques open exciting possibilities for models excelling in diverse tasks involving text, audio, and image processing. Models like GPT-4V, blending computer vision and language modeling, excel in complex…

Computation and Language · Computer Science 2023-10-20 Xiang Zhang , Senyu Li , Zijun Wu , Ning Shi

Generative audio models are rapidly advancing in both capabilities and public utilization -- several powerful generative audio models have readily available open weights, and some tech companies have released high quality generative audio…

Audio-visual semantic segmentation (AVSS) aims to segment and classify sounding objects in videos with acoustic cues. However, most approaches operate on the close-set assumption and only identify pre-defined categories from training data,…

Multimedia · Computer Science 2024-08-01 Ruohao Guo , Liao Qu , Dantong Niu , Yanyu Qi , Wenzhen Yue , Ji Shi , Bowei Xing , Xianghua Ying

This paper offers a mini review of Visual Word Sense Disambiguation (VWSD), which is a multimodal extension of traditional Word Sense Disambiguation (WSD). VWSD helps tackle lexical ambiguity in vision-language tasks. While conventional WSD…

Computation and Language · Computer Science 2026-02-03 Shashini Nilukshi , Deshan Sumanathilaka

Anomaly recognition plays a vital role in surveillance, transportation, healthcare, and public safety. However, most existing approaches rely solely on visual data, making them unreliable under challenging conditions such as occlusion, low…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Amjid Ali , Zulfiqar Ahmad Khan , Altaf Hussain , Muhammad Munsif , Adnan Hussain , Sung Wook Baik

Training Transformer-based models demands a large amount of data, while obtaining aligned and labelled data in multimodality is rather cost-demanding, especially for audio-visual speech recognition (AVSR). Thus it makes a lot of sense to…

Sound · Computer Science 2022-03-29 Xichen Pan , Peiyu Chen , Yichen Gong , Helong Zhou , Xinbing Wang , Zhouhan Lin

Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a video) and multi-modal events (i.e., those occurring in…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Yung-Hsuan Lai , Janek Ebbers , Yu-Chiang Frank Wang , François Germain , Michael Jeffrey Jones , Moitreya Chatterjee

This survey provides a comprehensive overview of recent advances in multimodal alignment and fusion within the field of machine learning, driven by the increasing availability and diversity of data modalities such as text, images, audio,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Songtao Li , Hao Tang

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic videos are disregarded.…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Jiawei Liu , Weining Wang , Sihan Chen , Xinxin Zhu , Jing Liu

Multimodal pathological images are usually in clinical diagnosis, but computer vision-based multimodal image-assisted diagnosis faces challenges with modality fusion, especially in the absence of expert-annotated data. To achieve the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Qinghua Lin , Guang-Hai Liu , Zuoyong Li , Yang Li , Yuting Jiang , Xiang Wu

Recent studies have identified that language models, pretrained on text-only datasets, often lack elementary visual knowledge, \textit{e.g.,} colors of everyday objects. Motivated by this observation, we ask whether a similar shortcoming…

Computation and Language · Computer Science 2025-01-17 Hyunjong Ok , Suho Yoo , Jaeho Lee

Recent advances in pre-trained vision transformers have shown promise in parameter-efficient audio-visual learning without audio pre-training. However, few studies have investigated effective methods for aligning multimodal features in…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Tanvir Mahmud , Shentong Mo , Yapeng Tian , Diana Marculescu

In this paper, we introduce a new problem, named audio-visual video parsing, which aims to parse a video into temporal event segments and label them as either audible, visible, or both. Such a problem is essential for a complete…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Yapeng Tian , Dingzeyu Li , Chenliang Xu

Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio,…

Multimedia · Computer Science 2026-05-15 Jianghan Chao , Jianzhang Gao , Wenhui Tan , Yuchong Sun , Ruihua Song , Liyun Ru

The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data interpretation. These models combine various data modalities,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Ziyue Huang , Hongxi Yan , Qiqi Zhan , Shuai Yang , Mingming Zhang , Chenkai Zhang , YiMing Lei , Zeming Liu , Qingjie Liu , Yunhong Wang
‹ Prev 1 8 9 10 Next ›