English

Domain Agnostic Few-shot Learning for Speaker Verification

Sound 2022-06-29 v1 Machine Learning Audio and Speech Processing

Abstract

Deep learning models for verification systems often fail to generalize to new users and new environments, even though they learn highly discriminative features. To address this problem, we propose a few-shot domain generalization framework that learns to tackle distribution shift for new users and new domains. Our framework consists of domain-specific and domain-aggregation networks, which are the experts on specific and combined domains, respectively. By using these networks, we generate episodes that mimic the presence of both novel users and novel domains in the training phase to eventually produce better generalization. To save memory, we reduce the number of domain-specific networks by clustering similar domains together. Upon extensive evaluation on artificially generated noise domains, we can explicitly show generalization ability of our framework. In addition, we apply our proposed methods to the existing competitive architecture on the standard benchmark, which shows further performance improvements.

Keywords

Cite

@article{arxiv.2206.13700,
  title  = {Domain Agnostic Few-shot Learning for Speaker Verification},
  author = {Seunghan Yang and Debasmit Das and Janghoon Cho and Hyoungwoo Park and Sungrack Yun},
  journal= {arXiv preprint arXiv:2206.13700},
  year   = {2022}
}

Comments

Proceedings of INTERSPEECH 2022

R2 v1 2026-06-24T12:06:14.317Z