中文
相关论文

相关论文: Learning Audio-Visual embedding for Person Verific…

200 篇论文

We propose an approach to extract speaker embeddings that are robust to speaking style variations in text-independent speaker verification. Typically, speaker embedding extraction includes training a DNN for speaker classification and using…

音频与语音处理 · 电气工程与系统科学 2022-06-29 Amber Afshan , Abeer Alwan

This study investigates the use of ear as a biometric for authentication and shows experimental results obtained on a newly created dataset of 420 images. Images are passed to a quality module in order to reduce False Rejection Rate. The…

密码学与安全 · 计算机科学 2009-12-08 Nazmeen Bibi Boodoo , R. K. Subramanian

The novelty of this study consists in a multi-modality approach to scene classification, where image and audio complement each other in a process of deep late fusion. The approach is demonstrated on a difficult classification problem,…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Jordan J. Bird , Diego R. Faria , Cristiano Premebida , Anikó Ekárt , George Vogiatzis

Robust audio-visual speech recognition (AVSR) in noisy environments remains challenging, as existing systems struggle to estimate audio reliability and dynamically adjust modality reliance. We propose router-gated cross-modal feature…

计算机视觉与模式识别 · 计算机科学 2025-08-27 DongHoon Lim , YoungChae Kim , Dong-Hyun Kim , Da-Hee Yang , Joon-Hyuk Chang

Verifying the identity of a speaker is crucial in modern human-machine interfaces, e.g., to ensure privacy protection or to enable biometric authentication. Classical speaker verification (SV) approaches estimate a fixed-dimensional…

音频与语音处理 · 电气工程与系统科学 2022-06-29 Ahmad Aloradi , Wolfgang Mack , Mohamed Elminshawi , Emanuël A. P. Habets

Embedding audio signal segments into vectors with fixed dimensionality is attractive because all following processing will be easier and more efficient, for example modeling, classifying or indexing. Audio Word2Vec previously proposed was…

计算与语言 · 计算机科学 2018-11-08 Sung-Feng Huang , Yi-Chen Chen , Hung-yi Lee , Lin-shan Lee

In this paper we propose a multi-modal multi-correlation learning framework targeting at the task of audio-visual speech separation. Although previous efforts have been extensively put on combining audio and visual modalities, most of them…

声音 · 计算机科学 2022-07-05 Xiaoyu Wang , Xiangyu Kong , Xiulian Peng , Yan Lu

Speaker embeddings extracted with deep 2D convolutional neural networks are typically modeled as projections of first and second order statistics of channel-frequency pairs onto a linear layer, using either average or attentive pooling…

音频与语音处理 · 电气工程与系统科学 2021-07-08 Themos Stafylakis , Johan Rohdin , Lukas Burget

In this work, we propose a new pooling strategy for language identification by considering Indian languages. The idea is to obtain utterance level features for any variable length audio for robust language recognition. We use the GhostVLAD…

计算与语言 · 计算机科学 2020-02-06 Krishna D N , Ankita Patil , M. S. P Raj , Sai Prasad H S , Prabhu Aashish Garapati

In this paper we explore audiovisual emotion recognition under noisy acoustic conditions with a focus on speech features. We attempt to answer the following research questions: (i) How does speech emotion recognition perform on noisy data?…

声音 · 计算机科学 2021-03-03 Michael Neumann , Ngoc Thang Vu

Audio-visual learning has been a major pillar of multi-modal machine learning, where the community mostly focused on its modality-aligned setting, i.e., the audio and visual modality are both assumed to signal the prediction target. With…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Yung-Hsuan Lai , Yen-Chun Chen , Yu-Chiang Frank Wang

This work introduces VERSE, a methodology for analyzing and improving Vision-Language Models applied to Visually-rich Document Understanding by exploring their visual embedding space. VERSE enables the visualization of latent…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Ignacio de Rodrigo , Alvaro J. Lopez-Lopez , Jaime Boal

This paper presents the results of the SUN team for the Compound Expressions Recognition Challenge of the 6th ABAW Competition. We propose a novel audio-visual method for compound expression recognition. Our method relies on emotion…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Elena Ryumina , Maxim Markitantov , Dmitry Ryumin , Heysem Kaya , Alexey Karpov

This paper presents Team Xaiofei's innovative approach to exploring Face-Voice Association in Multilingual Environments (FAME) at ACM Multimedia 2024. We focus on the impact of different languages in face-voice matching by building upon…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Jiehui Tang , Xiaofei Wang , Zhen Xiao , Jiayi Liu , Xueliang Liu , Richang Hong

Visual kinship verification entails confirming whether or not two individuals in a given pair of images or videos share a hypothesized kin relation. As a generalized face verification task, visual kinship verification is particularly…

计算机视觉与模式识别 · 计算机科学 2019-06-25 Xiaoting Wu , Eric Granger , Xiaoyi Feng

In this paper, we introduce a new problem, named audio-visual video parsing, which aims to parse a video into temporal event segments and label them as either audible, visible, or both. Such a problem is essential for a complete…

计算机视觉与模式识别 · 计算机科学 2020-07-23 Yapeng Tian , Dingzeyu Li , Chenliang Xu

Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using only coarse-grained labels. Faced with the challenging task settings, existing research…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Huilai Li , Xiaomeng Di , Ying Xing , Yonghao Dang , Yiming Wang , Jianqin Yin

Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…

音频与语音处理 · 电气工程与系统科学 2024-09-10 Hongyu Wang , Hui Li , Bo Li

In audiovisual automatic speech recognition (AV-ASR) systems, information fusion of visual features in a pre-trained ASR has been proven as a promising method to improve noise robustness. In this work, based on the prominent Whisper ASR,…

音频与语音处理 · 电气工程与系统科学 2026-01-27 Zhengyang Li , Thomas Graave , Björn Möller , Zehang Wu , Matthias Franz , Tim Fingscheidt

Audio-visual large language models (LLM) have drawn significant attention, yet the fine-grained combination of both input streams is rather under-explored, which is challenging but necessary for LLMs to understand general video inputs. To…

音频与语音处理 · 电气工程与系统科学 2023-10-11 Guangzhi Sun , Wenyi Yu , Changli Tang , Xianzhao Chen , Tian Tan , Wei Li , Lu Lu , Zejun Ma , Chao Zhang