English

Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech

Audio and Speech Processing 2026-01-21 v1

Abstract

Personal Voice Activity Detection (PVAD) is crucial for identifying target speaker segments in the mixture, yet its performance heavily depends on the quality of speaker embeddings. A key practical limitation is the short enrollment speech--such as a wake-up word--which provides limited cues. This paper proposes a novel adaptive speaker embedding self-augmentation strategy that enhances PVAD performance by augmenting the original enrollment embeddings through additive fusion of keyframe embeddings extracted from mixed speech. Furthermore, we introduce a long-term adaptation strategy to iteratively refine embeddings during detection, mitigating speaker temporal variability. Experiments show significant gains in recall, precision, and F1-score under short enrollment conditions, matching full-length enrollment performance after five iterative updates. The source code is available at https://anonymous.4open.science/r/ASE-PVAD-E5D6 .

Keywords

Cite

@article{arxiv.2601.12769,
  title  = {Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech},
  author = {Fuyuan Feng and Wenbin Zhang and Yu Gao and Longting Xu and Xiaofeng Mou and Yi Xu},
  journal= {arXiv preprint arXiv:2601.12769},
  year   = {2026}
}

Comments

Accepted by ICASSP 2026

R2 v1 2026-07-01T09:10:07.353Z