English

Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation

Sound 2025-02-06 v1 Artificial Intelligence Computation and Language Audio and Speech Processing

Abstract

Streaming multi-talker speech translation is a task that involves not only generating accurate and fluent translations with low latency but also recognizing when a speaker change occurs and what the speaker's gender is. Speaker change information can be used to create audio prompts for a zero-shot text-to-speech system, and gender can help to select speaker profiles in a conventional text-to-speech model. We propose to tackle streaming speaker change detection and gender classification by incorporating speaker embeddings into a transducer-based streaming end-to-end speech translation model. Our experiments demonstrate that the proposed methods can achieve high accuracy for both speaker change detection and gender classification.

Keywords

Cite

@article{arxiv.2502.02683,
  title  = {Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation},
  author = {Peidong Wang and Naoyuki Kanda and Jian Xue and Jinyu Li and Xiaofei Wang and Aswin Shanmugam Subramanian and Junkun Chen and Sunit Sivasankaran and Xiong Xiao and Yong Zhao},
  journal= {arXiv preprint arXiv:2502.02683},
  year   = {2025}
}
R2 v1 2026-06-28T21:32:41.448Z