中文

PAEFF:面向面部-语音关联的精准对齐与增强门控特征融合

计算机视觉与模式识别 2025-05-29 v2 人工智能

摘要

我们研究了学习面部与语音关联的任务,这最近在多模态社区中受到关注。这些方法受制于对负采样程序的刻意设计以及对 distant margin 参数的依赖。 These issues are addressed by learning a joint embedding space in which orthogonality constraints are applied to the fused embeddings of faces and voices. However, embedding spaces of faces and voices possess different characteristics and require spaces to be aligned before fusing them. To this end, we propose a method that accurately aligns the embedding spaces and fuses them with an enhanced gated fusion thereby improving the performance of face-voice association. Extensive experiments on the VoxCeleb dataset reveals the merits of the proposed approach.

关键词

引用

@article{arxiv.2505.17002,
  title  = {PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association},
  author = {Abdul Hannan and Muhammad Arslan Manzoor and Shah Nawaz and Muhammad Irzam Liaqat and Markus Schedl and Mubashir Noman},
  journal= {arXiv preprint arXiv:2505.17002},
  year   = {2025}
}

备注

Accepted at InterSpeech 2025