English

DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining

Audio and Speech Processing 2026-03-10 v1 Computation and Language Sound

Abstract

Speech-to-speech models handle turn-taking naturally but offer limited support for tool-calling or complex reasoning, while production ASR-LLM-TTS voice pipelines offer these capabilities but rely on silence timeouts, which lead to unnatural turn-taking. We present DualTurn, which narrows this gap through generative pretraining on dual-channel conversational audio. The model generates both speakers' future audio autoregressively, implicitly learning conversational dynamics without any labels, and is then fine-tuned to predict interpretable turn-taking signals that map directly to agent actions. DualTurn monitors both channels continuously, anticipating turn boundaries and producing five agent actions. On standard benchmarks, DualTurn (0.5B) outperforms both VAP on agent action prediction (wF1 0.633 vs. 0.389) and a 3.1B audio-text model on word-level turn prediction (AUC 0.930 vs. 0.880), while anticipating turn boundaries earlier with fewer interruptions.

Keywords

Cite

@article{arxiv.2603.08216,
  title  = {DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining},
  author = {Shangeth Rajaa},
  journal= {arXiv preprint arXiv:2603.08216},
  year   = {2026}
}

Comments

Submitted to Interspeech 2026

R2 v1 2026-07-01T11:10:03.672Z