English
Related papers

Related papers: DuplexCascade: Full-Duplex Speech-to-Speech Dialog…

200 papers

Full-duplex voice interaction allows users and agents to speak simultaneously with controllable barge-in, enabling lifelike assistants and customer service. Existing solutions are either end-to-end, difficult to design and hard to control,…

Achieving full-duplex communication in spoken dialogue systems (SDS) requires real-time coordination between listening, speaking, and thinking. This paper proposes a semantic voice activity detection (VAD) module as a dialogue manager (DM)…

Computation and Language · Computer Science 2025-02-26 Hao Zhang , Weiwei Li , Rilin Chen , Vinay Kothapally , Meng Yu , Dong Yu

Most end-to-end (E2E) spoken dialogue systems (SDS) rely on voice activity detection (VAD) for turn-taking, but VAD fails to distinguish between pauses and turn completions. Duplex SDS models address this by predicting output continuously,…

Computation and Language · Computer Science 2025-10-03 Siddhant Arora , Jinchuan Tian , Hayato Futami , Jiatong Shi , Yosuke Kashiwagi , Emiru Tsunoo , Shinji Watanabe

Achieving human-like responsiveness is a critical yet challenging goal for cascaded spoken dialogue systems. Conventional ASR-LLM-TTS pipelines follow a strictly sequential paradigm, requiring complete transcription and full reasoning…

Computation and Language · Computer Science 2026-02-27 Siyuan Liu , Jiahui Xu , Feng Jiang , Kuang Wang , Zefeng Zhao , Chu-Ren Huang , Jinghang Gu , Changqing Yin , Haizhou Li

Recent advances in spoken dialogue language models have shifted from turn-based to full-duplex designs, where the model continuously listens to the user while generating responses. However, existing duplex backbones still lack a native…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Haoyang Zhang , Jun Chen , Donghang Wu , Yuxin Li , Yuxin Zhang , Xiangyu Tony Zhang , Che Liu , Qingjian Lin , Yizhou Peng , Hexin Liu , Eng Siong Chng , Chao Yan , Boyong Wu , Yechang Huang , Xuerui Yang , Fei Tian

We present X-Talk, an open-source framework that champions a decoupled, modular design for LLM-driven speech-to-speech (S2S) systems. While the dominant trend favors end-to-end (E2E) modeling to optimize information flow, these…

Full-duplex speech interaction, as the most natural and intuitive mode of human communication, is driving artificial intelligence toward more human-like conversational systems. Traditional cascaded speech processing pipelines suffer from…

Artificial Intelligence · Computer Science 2026-05-01 Yadong Li , Guoxin Wu , Haiping Hou , Biye Li

Recent advances in AudioLLMs have enabled spoken dialogue systems to move beyond turn-based interaction toward real-time full-duplex communication, where the agent must decide when to speak, yield, or interrupt while the user is still…

Full-Duplex Speech Language Models (FD-SLMs) enable real-time, overlapping conversational interactions, offering a more dynamic user experience compared to traditional half-duplex models. However, existing benchmarks primarily focus on…

Computation and Language · Computer Science 2026-04-20 He Zhang , Wenqian Cui , Haoning Xu , Xiaohui Li , Lei Zhu , Haoli Bai , Shaohua Ma , Irwin King

Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumulate. In this paper, we propose a transducer-based speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-07 Jinzheng Zhao , Niko Moritz , Egor Lakomkin , Ruiming Xie , Zhiping Xiu , Katerina Zmolikova , Zeeshan Ahmed , Yashesh Gaur , Duc Le , Christian Fuegen

The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Tatiana Likhomanenko , Luke Carlson , Richard He Bai , Zijin Gu , Han Tran , Zakaria Aldeneh , Yizhe Zhang , Ruixiang Zhang , Huangjie Zheng , Navdeep Jaitly

Real-time speech-to-speech (S2S) models excel at generating natural, low-latency conversational responses but often lack deep knowledge and semantic understanding. Conversely, cascaded systems combining automatic speech recognition, a…

Computation and Language · Computer Science 2026-05-26 So Kuroki , Yotaro Kubo , Takuya Akiba , Yujin Tang

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and…

Computation and Language · Computer Science 2025-01-06 Qinglin Zhang , Luyao Cheng , Chong Deng , Qian Chen , Wen Wang , Siqi Zheng , Jiaqing Liu , Hai Yu , Chaohong Tan , Zhihao Du , Shiliang Zhang

True Full-Duplex (TFD) voice communication--enabling simultaneous listening and speaking with natural turn-taking, overlapping speech, and interruptions--represents a critical milestone toward human-like AI interaction. This survey…

Computation and Language · Computer Science 2025-09-19 Yuxuan Chen , Haoyuan Yu

Real-time spoken dialogue systems face a fundamental tension between latency and response quality. End-to-end speech-to-speech (S2S) models respond immediately and naturally handle turn-taking, backchanneling, and interruption, but produce…

Artificial Intelligence · Computer Science 2026-03-25 Long Mai

Recent advancements in large language models (LLMs) have led to significant progress in text-based dialogue systems. These systems can now generate high-quality responses that are accurate and coherent across a wide range of topics and…

Computation and Language · Computer Science 2025-01-10 Long Mai , Julie Carson-Berndsen

As large language models (LLMs) increasingly permeate daily lives, there is a growing demand for real-time interactions that mirror human conversations. Traditional turn-based chat systems driven by LLMs prevent users from verbally…

Computation and Language · Computer Science 2026-01-14 Xinrong Zhang , Yingfa Chen , Shengding Hu , Xu Han , Zihang Xu , Yuanwei Xu , Weilin Zhao , Maosong Sun , Zhiyuan Liu

Real-time voice agents face a dilemma: end-to-end models often lack deep reasoning, while cascaded pipelines incur high latency by executing ASR, LLM reasoning, and TTS strictly in sequence, unlike human conversation where listeners often…

Sound · Computer Science 2026-01-29 Wenhao Zou , Yuwei Miao , Zhanyu Ma , Jun Xu , Jiuchong Gao , Jinghua Hao , Renqing He , Jingwen Xu

As the paradigm of AI shifts from text-based LLMs to Speech Language Models (SLMs), there is a growing demand for full-duplex systems capable of real-time, natural human-computer interaction. However, the development of such models is…

Sound · Computer Science 2026-03-31 Kyudan Jung , Jihwan Kim , Soyoon Kim , Jeonghoon Kim , Jaegul Choo , Cheonbok Park

Full-duplex voice interaction is crucial for natural human computer interaction. We present a framework that decomposes complex dialogue into minimal conversational units, enabling the system to process each unit independently and predict…

Computation and Language · Computer Science 2026-01-30 Haoyuan Yu , Yuxuan Chen , Minjie Cai
‹ Prev 1 2 3 10 Next ›