English
Related papers

Related papers: MultiTurnCleanup: A Benchmark for Multi-Turn Spoke…

200 papers

In multilingual societies, social conversations often involve code-mixed speech. The current speech technology may not be well equipped to extract information from multi-lingual multi-speaker conversations. The DISPLACE challenge entails a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-06 Shikha Baghel , Shreyas Ramoji , Sidharth , Ranjana H , Prachi Singh , Somil Jain , Pratik Roy Chowdhuri , Kaustubh Kulkarni , Swapnil Padhi , Deepu Vijayasenan , Sriram Ganapathy

Language technologies should be judged on their usefulness in real-world use cases. An often overlooked aspect in natural language processing (NLP) research and evaluation is language variation in the form of non-standard dialects or…

Computation and Language · Computer Science 2024-07-09 Fahim Faisal , Orevaoghene Ahia , Aarohi Srivastava , Kabir Ahuja , David Chiang , Yulia Tsvetkov , Antonios Anastasopoulos

Accurate multi-turn intent classification is essential for advancing conversational AI systems. However, challenges such as the scarcity of comprehensive datasets and the complexity of contextual dependencies across dialogue turns hinder…

Computation and Language · Computer Science 2024-11-20 Junhua Liu , Yong Keat Tan , Bin Fu , Kwan Hui Lim

Task-oriented dialogue (TOD) systems are commonly designed with the presumption that each utterance represents a single intent. However, this assumption may not accurately reflect real-world situations, where users frequently express…

Computation and Language · Computer Science 2024-03-28 Yejin Yoon , Jungyeon Lee , Kangsan Kim , Chanhee Park , Taeuk Kim

Detecting and segmenting dysfluencies is crucial for effective speech therapy and real-time feedback. However, most methods only classify dysfluencies at the utterance level. We introduce StutterCut, a semi-supervised framework that…

Sound · Computer Science 2025-08-05 Suhita Ghosh , Melanie Jouaiti , Jan-Ole Perschewski , Sebastian Stober

Conversational tones -- the manners and attitudes in which speakers communicate -- are essential to effective communication. Amidst the increasing popularization of Large Language Models (LLMs) over recent years, it becomes necessary to…

Computation and Language · Computer Science 2024-06-07 Dun-Ming Huang , Pol Van Rijn , Ilia Sucholutsky , Raja Marjieh , Nori Jacoby

Neural network based speech dereverberation has achieved promising results in recent studies. Nevertheless, many are focused on recovery of only the direct path sound and early reflections, which could be beneficial to speech perception,…

Sound · Computer Science 2021-10-19 Ziteng Wang , Yueyue Na , Biao Tian , Qiang Fu

Evaluating disfluency removal in speech requires more than aggregate token-level scores. Traditional word-based metrics such as precision, recall, and F1 (E-Scores) capture overall performance but cannot reveal why models succeed or fail.…

Language models are increasingly deployed in interactive settings where users reason about facts over time rather than in isolation. In such scenarios, correct behavior requires models to maintain and update implicit temporal assumptions…

Computation and Language · Computer Science 2026-04-28 Yash Kumar Atri , Steven L. Johnson , Tom Hartvigsen

Multiturn dialogue models aim to generate human-like responses by leveraging conversational context, consisting of utterances from previous exchanges. Existing methods often neglect the interactions between these utterances or treat all of…

Computation and Language · Computer Science 2025-04-15 Akanksha Mehndiratta , Krishna Asawa

Full-duplex interaction is crucial for natural human-machine communication, yet remains challenging as it requires robust turn-taking detection to decide when the system should speak, listen, or remain silent. Existing solutions either rely…

Computation and Language · Computer Science 2025-09-30 Guojian Li , Chengyou Wang , Hongfei Xue , Shuiyuan Wang , Dehui Gao , Zihan Zhang , Yuke Lin , Wenjie Li , Longshuai Xiao , Zhonghua Fu , Lei Xie

Recently, the source separation performance was greatly improved by time-domain audio source separation based on dual-path recurrent neural network (DPRNN). DPRNN is a simple but effective model for a long sequential data. While DPRNN is…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-25 Keisuke Kinoshita , Thilo von Neumann , Marc Delcroix , Tomohiro Nakatani , Reinhold Haeb-Umbach

While NLP research has made strides in conversational tasks, many approaches focus on single-turn responses with well-defined objectives or evaluation criteria. In contrast, coaching presents unique challenges with initially undefined goals…

Computation and Language · Computer Science 2025-07-10 Vidya Srinivas , Xuhai Xu , Xin Liu , Kumar Ayush , Isaac Galatzer-Levy , Shwetak Patel , Daniel McDuff , Tim Althoff

This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets, we propose an…

Computation and Language · Computer Science 2025-05-21 Yuxin Lin , Yinglin Zheng , Ming Zeng , Wangzheng Shi

Dialogue related Machine Reading Comprehension requires language models to effectively decouple and model multi-turn dialogue passages. As a dialogue development goes after the intentions of participants, its topic may not keep constant…

Computation and Language · Computer Science 2023-09-19 Xinbei Ma , Yi Xu , Hai Zhao , Zhuosheng Zhang

Providing voice assistants the ability to navigate multi-turn conversations is a challenging problem. Handling multi-turn interactions requires the system to understand various conversational use-cases, such as steering, intent carryover,…

Computation and Language · Computer Science 2023-10-30 Jiarui Lu , Bo-Hsiang Tseng , Joel Ruben Antony Moniz , Site Li , Xueyun Zhu , Hong Yu , Murat Akbacak

The rapid advancement of talking-head deepfake generation fueled by advanced generative models has elevated the realism of synthetic videos to a level that poses substantial risks in domains such as media, politics, and finance. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Xinqi Xiong , Prakrut Patel , Qingyuan Fan , Amisha Wadhwa , Sarathy Selvam , Xiao Guo , Luchao Qi , Xiaoming Liu , Roni Sengupta

Speaker counting is the task of estimating the number of people that are simultaneously speaking in an audio recording. For several audio processing tasks such as speaker diarization, separation, localization and tracking, knowing the…

Sound · Computer Science 2021-01-07 Pierre-Amaury Grumiaux , Srdan Kitic , Laurent Girin , Alexandre Guérin

In this paper, we propose a novel spoken-text-style conversion method that can simultaneously execute multiple style conversion modules such as punctuation restoration and disfluency deletion without preparing matched datasets. In practice,…

Computation and Language · Computer Science 2021-06-24 Mana Ihori , Naoki Makishima , Tomohiro Tanaka , Akihiko Takashima , Shota Orihashi , Ryo Masumura

Speech disfluency commonly occurs in conversational and spontaneous speech. However, standard Automatic Speech Recognition (ASR) models struggle to accurately recognize these disfluencies because they are typically trained on fluent…

Computation and Language · Computer Science 2024-09-18 Robin Amann , Zhaolin Li , Barbara Bruno , Jan Niehues