English
Related papers

Related papers: FireRedChat: A Pluggable, Full-Duplex Voice Intera…

200 papers

Recent advances in duplex speech models have enabled natural, low-latency speech-to-speech interactions. However, existing models are restricted to a fixed role and voice, limiting their ability to support structured, role-driven real-world…

Computation and Language · Computer Science 2026-02-09 Rajarshi Roy , Jonathan Raiman , Sang-gil Lee , Teodor-Dumitru Ene , Robert Kirby , Sungwon Kim , Jaehyeon Kim , Bryan Catanzaro

Full-Duplex Speech Language Models (FD-SLMs) are specialized foundation models designed to enable natural, real-time spoken interactions by modeling complex conversational turn-taking such as interruptions, backchannels, and overlapping…

Computation and Language · Computer Science 2026-01-21 Wenqian Cui , Lei Zhu , Xiaohui Li , Zhihan Guo , Haoli Bai , Lu Hou , Irwin King

In Task-Oriented Dialogue (TOD) systems, correctly updating the system's understanding of the user's requests (\textit{a.k.a} dialogue state tracking) is key to a smooth interaction. Traditionally, TOD systems perform this update in three…

Computation and Language · Computer Science 2024-07-02 Lucas Druart , Valentin Vielzeuf , Yannick Estève

Recent advances in spoken dialogue systems have brought increased attention to human-like full-duplex voice interactions. However, our comprehensive review of this field reveals several challenges, including the difficulty in obtaining…

Immersive conversational systems in production face a persistent trade-off between responsiveness and long-horizon task capability. Real-time interaction is achievable for lightweight turns, but requests involving planning and tool…

Spoken dialogue is an intuitive form of human-computer interaction, yet current speech language models often remain constrained to turn-based exchanges, lacking real-time adaptability such as user barge-in. We propose a novel duplex speech…

Computation and Language · Computer Science 2025-07-28 Ke Hu , Ehsan Hosseini-Asl , Chen Chen , Edresson Casanova , Subhankar Ghosh , Piotr Żelasko , Zhehuai Chen , Jason Li , Jagadeesh Balam , Boris Ginsburg

The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Tatiana Likhomanenko , Luke Carlson , Richard He Bai , Zijin Gu , Han Tran , Zakaria Aldeneh , Yizhe Zhang , Ruixiang Zhang , Huangjie Zheng , Navdeep Jaitly

Spoken dialogue models have significantly advanced intelligent human-computer interaction, yet they lack a plug-and-play full-duplex prediction module for semantic endpoint detection, hindering seamless audio interactions. In this paper, we…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-05 Weijie Wu , Wenhao Guan , Kaidi Wang , Peijie Chen , Zhuanling Zha , Junbo Li , Jun Fang , Lin Li , Qingyang Hong

Full-Duplex Speech Language Models (FD-SLMs) enable real-time, overlapping conversational interactions, offering a more dynamic user experience compared to traditional half-duplex models. However, existing benchmarks primarily focus on…

Computation and Language · Computer Science 2026-04-20 He Zhang , Wenqian Cui , Haoning Xu , Xiaohui Li , Lei Zhu , Haoli Bai , Shaohua Ma , Irwin King

We present a generative dialogue system capable of operating in a full-duplex manner, allowing for seamless interaction. It is based on a large language model (LLM) carefully aligned to be aware of a perception module, a motor function…

Computation and Language · Computer Science 2024-10-30 Peng Wang , Songshuo Lu , Yaohua Tang , Sijie Yan , Wei Xia , Yuanjun Xiong

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and…

Computation and Language · Computer Science 2025-01-06 Qinglin Zhang , Luyao Cheng , Chong Deng , Qian Chen , Wen Wang , Siqi Zheng , Jiaqing Liu , Hai Yu , Chaohong Tan , Zhihao Du , Shiliang Zhang

Real-time speech conversation is essential for natural and efficient human-machine interactions, requiring duplex and streaming capabilities. Traditional Transformer-based conversational chatbots operate in a turn-based manner and exhibit…

Computation and Language · Computer Science 2025-04-04 Xiangyu Lu , Wang Xu , Haoyu Wang , Hongyun Zhou , Haiyan Zhao , Conghui Zhu , Tiejun Zhao , Muyun Yang

Real-time voice agents face a dilemma: end-to-end models often lack deep reasoning, while cascaded pipelines incur high latency by executing ASR, LLM reasoning, and TTS strictly in sequence, unlike human conversation where listeners often…

Sound · Computer Science 2026-01-29 Wenhao Zou , Yuwei Miao , Zhanyu Ma , Jun Xu , Jiuchong Gao , Jinghua Hao , Renqing He , Jingwen Xu

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

Computation and Language · Computer Science 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ziqiao Peng , Yanbo Fan , Haoyu Wu , Xuan Wang , Hongyan Liu , Jun He , Zhaoxin Fan

Full-duplex voice agents--systems that listen and speak simultaneously--are rapidly moving from research to production. However, existing evaluations address conversational dynamics and task completion in isolation. We introduce…

Sound · Computer Science 2026-03-17 Soham Ray , Keshav Dhandhania , Victor Barres , Karthik Narasimhan

Full-duplex dialog models aim to listen and speak simultaneously, delivering rapid responses to dynamic user input. Among different solutions to full-duplexity, a native solution merges multiple channels in each time step, achieving the…

Sound · Computer Science 2026-02-02 Yiqun Yao , Xiang Li , Xin Jiang , Xuezhi Fang , Naitong Yu , Wenjia Ma , Aixin Sun , Yequan Wang

Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smiles, and gestures. To support successful human-agent…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Amrita Mazumdar , Seonwook Park , Rajarshi Roy , Nikhil Srihari , Shengze Wang , Yuhao Zhou , Julia Wang , Koki Nagano , Shalini De Mello

As large language models (LLMs) increasingly permeate daily lives, there is a growing demand for real-time interactions that mirror human conversations. Traditional turn-based chat systems driven by LLMs prevent users from verbally…

Computation and Language · Computer Science 2026-01-14 Xinrong Zhang , Yingfa Chen , Shengding Hu , Xu Han , Zihang Xu , Yuanwei Xu , Weilin Zhao , Maosong Sun , Zhiyuan Liu

In this work, we propose a novel cross-talk rejection framework for a multi-channel multi-talker setup for a live multiparty interactive show. Our far-field audio setup is required to be hands-free during live interaction and comprises four…

Sound · Computer Science 2024-02-16 Hyewon Han , Naveen Kumar