面向多模态基础模型的语言无关人脸-语音关联
音频与语音处理
2025-12-03 v1 声音
图像与视频处理
摘要
本文描述了提交给FAME2026挑战赛的UZH-CL系统。该挑战聚焦于在独特多语言条件下的跨模态验证,具体指未见且未听过的语言。我们的ethods研究了两种不同的体系结构:一种是基于对比损失和正交投影损失训练的基线双编码器系统,另一种是利用ImageBind配合LoRA的基础模型方法。为解决该挑战的数据稀缺和语言限制问题,我们从VoxBlink处精选了外部阿拉伯语数据集。我们最佳的系统ImageBind-LoRA展现出显著的跨语言泛化能力:尽管仅在阿拉伯语音上进行微调,却在评估集(英文和德语)上实现了24.73%的等效错误率(EER),夺得赛事第二名。
引用
@article{arxiv.2512.02759,
title = {Towards Language-Independent Face-Voice Association with Multimodal Foundation Models},
author = {Aref Farhadipour and Teodora Vukovic and Volker Dellwo},
journal= {arXiv preprint arXiv:2512.02759},
year = {2025}
}
备注
This paper presents the system description of the UZH-CL team for the FAME2026 Challenge at ICASSP 2026. Our model achieved second place in the final ranking