English
Related papers

Related papers: Talking Turns: Benchmarking Audio Foundation Model…

200 papers

This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets, we propose an…

Computation and Language · Computer Science 2025-05-21 Yuxin Lin , Yinglin Zheng , Ming Zeng , Wangzheng Shi

Dialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM) have significantly enhanced speech-based conversational AI. However, these models are limited to turn-based…

Computation and Language · Computer Science 2024-08-06 Ziyang Ma , Yakun Song , Chenpeng Du , Jian Cong , Zhuo Chen , Yuping Wang , Yuxuan Wang , Xie Chen

Foundation models (FMs) are increasingly spearheading recent advances on a variety of tasks that fall under the purview of computer audition -- the use of machines to understand sounds. They feature several advantages over traditional…

Turn-taking is a fundamental aspect of human communication and can be described as the ability to take turns, project upcoming turn shifts, and supply backchannels at appropriate locations throughout a conversation. In this work, we…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-13 Erik Ekstedt , Gabriel Skantze

Predicting team dynamics from personality traits remains a fundamental challenge for the psychological sciences and team-based organizations. Understanding how team composition generates team processes can significantly advance team-based…

Computation and Language · Computer Science 2024-11-26 Lisa R. O'Bryan , Madeline Navarro , Juan Segundo Hevia , Santiago Segarra

Spoken Dialogue Models (SDMs) have advanced rapidly, yet their ability to sustain genuinely interactive multi-turn conversations remains underexplored, as most benchmarks focus on single-turn exchanges. We introduce Multi-Bench, the first…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-04 Yayue Deng , Guoqiang Hu , Haiyang Sun , Xiangyu Zhang , Haoyang Zhang , Fei Tian , Xuerui Yang , Gang Yu , Eng Siong Chng

Despite being trained exclusively on speech data, speech foundation models (SFMs) like Whisper have shown impressive performance in non-speech tasks such as audio classification. This is partly because speech shares some common traits with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-17 Orchid Chetia Phukan , Swarup Ranjan Behera , Girish , Mohd Mujtaba Akhtar , Arun Balaji Buduru , Rajesh Sharma

Turn-taking is a fundamental aspect of human communication where speakers convey their intention to either hold, or yield, their turn through prosodic cues. Using the recently proposed Voice Activity Projection model, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Erik Ekstedt , Siyang Wang , Éva Székely , Joakim Gustafson , Gabriel Skantze

Understanding the inner mechanisms of black-box foundation models (FMs) is essential yet challenging in artificial intelligence and its applications. Over the last decade, the long-running focus has been on their explainability, leading to…

Machine Learning · Computer Science 2024-11-26 Shi Fu , Yuzhu Chen , Yingjie Wang , Dacheng Tao

Speech-to-speech models handle turn-taking naturally but offer limited support for tool-calling or complex reasoning, while production ASR-LLM-TTS voice pipelines offer these capabilities but rely on silence timeouts, which lead to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Shangeth Rajaa

In human conversational interactions, turn-taking exchanges can be coordinated using cues from multiple modalities. To design spoken dialog systems that can conduct fluid interactions it is desirable to incorporate cues from separate…

Computation and Language · Computer Science 2018-09-03 Matthew Roddy , Gabriel Skantze , Naomi Harte

Open-domain dialog systems (also known as chatbots) have increasingly drawn attention in natural language processing. Some of the recent work aims at incorporating affect information into sequence-to-sequence neural dialog modeling, making…

Computation and Language · Computer Science 2020-06-25 Yubo Xie , Ekaterina Svikhnushina , Pearl Pu

Full-duplex spoken dialogue systems promise to transform human-machine interaction from a rigid, turn-based protocol into a fluid, natural conversation. However, the central challenge to realizing this vision, managing overlapping speech,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-28 Guan-Ting Lin , Shih-Yun Shan Kuan , Qirui Wang , Jiachen Lian , Tingle Li , Shinji Watanabe , Hung-yi Lee

Conversation is a subject of increasing interest in the social, cognitive, and computational sciences. Yet as conversational datasets continue to increase in size and complexity, researchers lack scalable methods to segment speech-to-text…

Computation and Language · Computer Science 2025-11-13 Gus Cooney , Andrew Reece

Machine Learning (ML) models are increasingly used to make critical decisions in real-world applications, yet they have become more complex, making them harder to understand. To this end, researchers have proposed several techniques to…

Machine Learning · Computer Science 2023-03-07 Dylan Slack , Satyapriya Krishna , Himabindu Lakkaraju , Sameer Singh

Existing voice AI assistants treat every detected pause as an invitation to speak. This works in dyadic dialogue, but in multi-party settings, where an AI assistant participates alongside multiple speakers, pauses are abundant and…

Artificial Intelligence · Computer Science 2026-03-13 Kratika Bhagtani , Mrinal Anand , Yu Chen Xu , Amit Kumar Singh Yadav

Despite their success in numerous fields, the potential of foundation models for modeling and understanding human behavior remains largely unexplored. We introduce Be.FM, one of the first open foundation models designed for human behavior…

Despite the multi-turn open-domain dialogue systems have attracted more and more attention and made great progress, the existing dialogue systems are still very boring. Nearly all the existing dialogue models only provide a response when…

Computation and Language · Computer Science 2019-12-23 Tian Lan , Xianling Mao , Heyan Huang , Wei Wei

We present a real-time front-end for voice-based conversational AI to enable natural turn-taking in two-speaker scenarios by combining primary speaker segmentation with hierarchical End-of-Turn (EOT) detection. To operate robustly in…

Machine Learning · Computer Science 2026-03-17 Karim Helwani , Hoang Do , James Luan , Sriram Srinivasan

Dialogue Act (DA) classification is the task of classifying utterances with respect to the function they serve in a dialogue. Existing approaches to DA classification model utterances without incorporating the turn changes among speakers…

Computation and Language · Computer Science 2021-09-14 Zihao He , Leili Tavabi , Kristina Lerman , Mohammad Soleymani