English

When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR

Computation and Language 2026-05-28 v1

Abstract

SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts with sensitive information, fine-tune on proprietary recordings, or both. We identify and systematically investigate an overlooked privacy risk of such customisation: a model adapted to recognise domain-specific terminology can be nudged into transcribing a phonetically similar word from its context or training data, even when a different word is spoken, thereby leaking private information. To evaluate this risk, we construct a controlled dataset and measure leakage rates across two customisation mechanisms, prompting and fine-tuning. Both mechanisms cause measurable leakage, compounding when combined. We evaluate a prompt-level mitigation strategy and analyse the accuracy-leakage trade-off across customisation approaches, finding that fine-tuning without context prompts offers the best balance. We release our code and dataset publicly.

Keywords

Cite

@article{arxiv.2605.28211,
  title  = {When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR},
  author = {Maike Züfle and Jan Niehues},
  journal= {arXiv preprint arXiv:2605.28211},
  year   = {2026}
}