中文

多模态适应与泛化的最新进展:从传统方法到基础模型

计算机视觉与模式识别 2025-09-22 v4 人工智能 机器学习 机器人学

摘要

在现实场景中,实现 domain adaptation 和 generalization 面临着重大挑战,因为模型必须适应或泛化到未知的目标分布。将这些能力扩展到未见的多模态分布(即多模态 domain adaptation 和 generalization)更加具有挑战性,因为不同模态具有各自的特征。多年来已取得显著进展,应用范围涵盖动作识别、语义分割等。此外,近期大规模预训练多模态基础模型的出现(如 CLIP),激发了利用这些模型来增强适应和泛化性能或将其适用于下游任务的工作。本综述提供了从传统方法到基础模型的最新进展的首个全面综述,涵盖:(1) 多模态 domain adaptation;(2) 多模态 test-time adaptation;(3) 多模态 domain generalization;(4) 利用多模态基础模型进行 domain adaptation 和 generalization;(5) 多模态基础模型的适应。对于每个主题,我们对问题进行形式化定义,并详尽综述现有方法。此外,我们分析了相关数据集和应用,突出了开放性挑战和潜在的未来研究方向。我们维护了一个包含最新文献的活跃存储库,地址为 https://github.com/donghao51/Awesome-Multimodal-Adaptation。

关键词

引用

@article{arxiv.2501.18592,
  title  = {Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models},
  author = {Hao Dong and Moru Liu and Kaiyang Zhou and Eleni Chatzi and Juho Kannala and Cyrill Stachniss and Olga Fink},
  journal= {arXiv preprint arXiv:2501.18592},
  year   = {2025}
}

备注

Project page: https://github.com/donghao51/Awesome-Multimodal-Adaptation