English
Related papers

Related papers: TOLD: A Novel Two-Stage Overlap-Aware Framework fo…

200 papers

We propose a data cleansing method that utilizes a neural analysis and synthesis (NANSY++) framework to train an end-to-end neural diarization model (EEND) for singer diarization. Our proposed model converts song data with choral singing…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-25 Hokuto Munakata , Ryo Terashima , Yusuke Fujita

Recently, end-to-end models have become a popular approach as an alternative to traditional hybrid models in automatic speech recognition (ASR). The multi-speaker speech separation and recognition task is a central task in cocktail party…

Computation and Language · Computer Science 2018-11-07 Xuankai Chang , Yanmin Qian , Kai Yu , Shinji Watanabe

Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-21 Ming Cheng , Weiqing Wang , Yucong Zhang , Xiaoyi Qin , Ming Li

Vision-language models (VLMs) such as CLIP exhibit strong Out-of-distribution (OOD) detection capabilities by aligning visual and textual representations. Recent CLIP-based test-time adaptation methods further improve detection performance…

Computation and Language · Computer Science 2026-04-20 Jinlun Ye , Jiang Liao , Runhe Lai , Xinhua Lu , Jiaxin Zhuang , Zhiyong Gan , Ruixuan Wang

Multi-talker overlapped speech poses a significant challenge for speech recognition and diarization. Recent research indicated that these two tasks are inter-dependent and complementary, motivating us to explore a unified modeling method to…

Sound · Computer Science 2023-05-26 Lingwei Meng , Jiawen Kang , Mingyu Cui , Haibin Wu , Xixin Wu , Helen Meng

This study investigates whether speech-based depression detection models learn depression-related acoustic biomarkers or instead rely on speaker identity cues. Using the DAIC-WOZ dataset, we propose a data-splitting strategy that controls…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-17 Hsiang-Chen Yeh , Luqi Sun , Aurosweta Mahapatra , Shreeram Suresh Chandra , Emily Mower Provost , Berrak Sisman

Serialized output training (SOT) attracts increasing attention due to its convenience and flexibility for multi-speaker automatic speech recognition (ASR). However, it is not easy to train with attention loss only. In this paper, we propose…

Sound · Computer Science 2024-09-12 Hao Shi , Yuan Gao , Zhaoheng Ni , Tatsuya Kawahara

Sound Event Detection (SED) detects regions of sound events, while Speaker Diarization (SD) segments speech conversations attributed to individual speakers. In SED, all speaker segments are classified as a single speech event, while in SD,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-16 Yidi Jiang , Ruijie Tao , Wen Huang , Qian Chen , Wen Wang

Transducer and Attention based Encoder-Decoder (AED) are two widely used frameworks for speech-to-text tasks. They are designed for different purposes and each has its own benefits and drawbacks for speech-to-text tasks. In order to…

Computation and Language · Computer Science 2023-05-08 Yun Tang , Anna Y. Sun , Hirofumi Inaguma , Xinyue Chen , Ning Dong , Xutai Ma , Paden D. Tomasello , Juan Pino

End-to-end architectures have been recently proposed for spoken language understanding (SLU) and semantic parsing. Based on a large amount of data, those models learn jointly acoustic and linguistic-sequential features. Such architectures…

Computation and Language · Computer Science 2020-02-17 Marco Dinarelli , Nikita Kapoor , Bassam Jabaian , Laurent Besacier

In this paper, different online speaker diarization systems are evaluated on the same hardware with the same test data with regard to their latency. The latency is the time span from audio input to the output of the corresponding speaker…

Computation and Language · Computer Science 2024-07-08 Roman Aperdannier , Sigurd Schacht , Alexander Piazza

The LEAP submission for DIHARD-III challenge is described in this paper. The proposed system is composed of a speech bandwidth classifier, and diarization systems fine-tuned for narrowband and wideband speech separately. We use an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-15 Prachi Singh , Rajat Varma , Venkat Krishnamohan , Srikanth Raj Chetupalli , Sriram Ganapathy

We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding…

Sound · Computer Science 2021-02-11 Zeqian Li , Jacob Whitehill

Large Language Models (LLMs) can be adapted to extend their text capabilities to speech inputs. However, these speech-adapted LLMs consistently underperform their text-based counterparts--and even cascaded pipelines--on language…

Computation and Language · Computer Science 2026-02-24 Santiago Cuervo , Skyler Seto , Maureen de Seyssel , Richard He Bai , Zijin Gu , Tatiana Likhomanenko , Navdeep Jaitly , Zakaria Aldeneh

Thus far, end-to-end (E2E) models have not been shown to outperform state-of-the-art conventional models with respect to both quality, i.e., word error rate (WER), and latency, i.e., the time the hypothesis is finalized after the user stops…

Speaker extraction and diarization are two enabling techniques for real-world speech applications. Speaker extraction aims to extract a target speaker's voice from a speech mixture, while speaker diarization demarcates speech segments by…

Sound · Computer Science 2025-01-17 Junyi Ao , Mehmet Sinan Yıldırım , Ruijie Tao , Meng Ge , Shuai Wang , Yanmin Qian , Haizhou Li

Semi-supervised domain adaptation (SSDA) methods have demonstrated great potential in large-scale image classification tasks when massive labeled data are available in the source domain but very few labeled samples are provided in the…

Computer Vision and Pattern Recognition · Computer Science 2020-12-07 Zhiyong Huang , Kekai Sheng , Weiming Dong , Xing Mei , Chongyang Ma , Feiyue Huang , Dengwen Zhou , Changsheng Xu

This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved…

Transcribing the speech of multiple overlapping speakers typically requires separating the audio into multiple streams and recognizing each one independently. More recent work jointly separates and transcribes, but requires a separate…

Computation and Language · Computer Science 2024-08-14 Chak-Fai Li , William Hartmann , Matthew Snover

End-to-end task-oriented dialogue (TOD) systems have achieved promising performance by leveraging sophisticated natural language understanding and natural language generation capabilities of pre-trained models. This work enables the TOD…

Computation and Language · Computer Science 2023-08-17 Jianguo Zhang , Stephen Roller , Kun Qian , Zhiwei Liu , Rui Meng , Shelby Heinecke , Huan Wang , Silvio Savarese , Caiming Xiong