中文
相关论文

相关论文: Dr. SHAP-AV: Decoding Relative Modality Contributi…

200 篇论文

SHAP is a popular method for measuring variable importance in machine learning models. In this paper, we study the algorithm used to estimate SHAP scores and outline its connection to the functional ANOVA decomposition. We use this…

统计方法学 · 统计学 2022-11-14 Andrew Herren , P. Richard Hahn

In audiovisual automatic speech recognition (AV-ASR) systems, information fusion of visual features in a pre-trained ASR has been proven as a promising method to improve noise robustness. In this work, based on the prominent Whisper ASR,…

音频与语音处理 · 电气工程与系统科学 2026-01-27 Zhengyang Li , Thomas Graave , Björn Möller , Zehang Wu , Matthias Franz , Tim Fingscheidt

Speaker-attributed automatic speech recognition (SA-ASR) aims to transcribe speech while assigning transcripts to the corresponding speakers accurately. Existing methods often rely on complex modular systems or require extensive fine-tuning…

计算与语言 · 计算机科学 2025-01-16 Thai-Binh Nguyen , Alexander Waibel

Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse questions about audio-visual scenes. However, existing methods lack sufficient flexibility and…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Jiayu Zhang , Shuo Ye , Qilang Ye , Xun Lin , Zihan Song , Zitong Yu

Recent advances in Audio-Visual Speech Recognition (AVSR) have led to unprecedented achievements in the field, improving the robustness of this type of system in adverse, noisy environments. In most cases, this task has been addressed…

计算机视觉与模式识别 · 计算机科学 2025-05-07 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Large language models (LLMs) have recently achieved impressive results in speech recognition across multiple modalities, including Auditory Speech Recognition (ASR), Visual Speech Recognition (VSR), and Audio-Visual Speech Recognition…

音频与语音处理 · 电气工程与系统科学 2026-01-28 Umberto Cappellazzo , Xubo Liu , Pingchuan Ma , Stavros Petridis , Maja Pantic

Audio-visual segmentation (AVS) is a challenging task that involves accurately segmenting sounding objects based on audio-visual cues. The effectiveness of audio-visual learning critically depends on achieving accurate cross-modal alignment…

计算机视觉与模式识别 · 计算机科学 2024-08-15 Yuanhong Chen , Yuyuan Liu , Hu Wang , Fengbei Liu , Chong Wang , Helen Frazer , Gustavo Carneiro

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

计算机视觉与模式识别 · 计算机科学 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

Audio-visual learning has been a major pillar of multi-modal machine learning, where the community mostly focused on its modality-aligned setting, i.e., the audio and visual modality are both assumed to signal the prediction target. With…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Yung-Hsuan Lai , Yen-Chun Chen , Yu-Chiang Frank Wang

Audio-visual recognition (AVR) has been considered as a solution for speech recognition tasks when the audio is corrupted, as well as a visual recognition method used for speaker verification in multi-speaker scenarios. The approach of AVR…

计算机视觉与模式识别 · 计算机科学 2017-11-01 Amirsina Torfi , Seyed Mehdi Iranmanesh , Nasser M. Nasrabadi , Jeremy Dawson

Robust audio-visual speech recognition (AVSR) in noisy environments remains challenging, as existing systems struggle to estimate audio reliability and dynamically adjust modality reliance. We propose router-gated cross-modal feature…

计算机视觉与模式识别 · 计算机科学 2025-08-27 DongHoon Lim , YoungChae Kim , Dong-Hyun Kim , Da-Hee Yang , Joon-Hyuk Chang

How to effectively interact audio with vision has garnered considerable interest within the multi-modality research field. Recently, a novel audio-visual segmentation (AVS) task has been proposed, aiming to segment the sounding objects in…

计算机视觉与模式识别 · 计算机科学 2024-02-07 Tianxiang Chen , Zhentao Tan , Tao Gong , Qi Chu , Yue Wu , Bin Liu , Le Lu , Jieping Ye , Nenghai Yu

Shapley values, a gold standard for feature attribution in Explainable AI, face two key challenges. First, the canonical Shapley framework assumes that the worth function is additive, yet real-world payoff constructions--driven by…

机器学习 · 计算机科学 2026-03-10 Jialai She

Modern automatic speech recognition (ASR) systems need to be robust under acoustic variability arising from environmental, speaker, channel, and recording conditions. Ensuring such robustness to variability is a challenge in modern day…

计算与语言 · 计算机科学 2016-12-07 Dmitriy Serdyuk , Kartik Audhkhasi , Philémon Brakel , Bhuvana Ramabhadran , Samuel Thomas , Yoshua Bengio

Recent Audio-Visual Question Answering (AVQA) methods have advanced significantly. However, most AVQA methods lack effective mechanisms for handling missing modalities, suffering from severe performance degradation in real-world scenarios…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Jiayu Zhang , Shuo Ye , Qilang Ye , Zihan Song , Jiajian Huang , Zitong Yu

Nowadays, we have witnessed the early progress on learning the association between voice and face automatically, which brings a new wave of studies to the computer vision community. However, most of the prior arts along this line (a) merely…

计算机视觉与模式识别 · 计算机科学 2021-03-15 Peisong Wen , Qianqian Xu , Yangbangyan Jiang , Zhiyong Yang , Yuan He , Qingming Huang

In this work, we propose an innovative framework that integrates EEG, image, and text data, aiming to decode visual neural representations from low signal-to-noise ratio EEG signals. Specifically, we introduce text modality to enhance the…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Kaili sun , Xingyu Miao , Bing Zhai , Haoran Duan , Yang Long

Audio-Visual Localization (AVL) aims to identify sound-emitting sources within a visual scene. However, existing studies focus on image-level audio-visual associations, failing to capture temporal dynamics. Moreover, they assume simplified…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Hahyeon Choi , Junhoo Lee , Nojun Kwak

Audiovisual automatic speech recognition (AV-ASR) aims to improve the robustness of a speech recognition system by incorporating visual information. Training fully supervised multimodal models for this task from scratch, however is limited…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Paul Hongsuck Seo , Arsha Nagrani , Cordelia Schmid

Audio-Visual Segmentation (AVS) aims to precisely outline audible objects in a visual scene at the pixel level. Existing AVS methods require fine-grained annotations of audio-mask pairs in supervised learning fashion. This limits their…

计算机视觉与模式识别 · 计算机科学 2023-09-14 Swapnil Bhosale , Haosen Yang , Diptesh Kanojia , Xiatian Zhu