中文
相关论文

相关论文: Overlap-aware diarization: resegmentation using ne…

200 篇论文

A novel framework for meeting transcription using asynchronous microphones is proposed in this paper. It consists of audio synchronization, speaker diarization, utterance-wise speech enhancement using guided source separation, automatic…

音频与语音处理 · 电气工程与系统科学 2020-08-03 Shota Horiguchi , Yusuke Fujita , Kenji Nagamatsu

While standard speaker diarization attempts to answer the question "who spoken when", most of relevant applications in reality are more interested in determining "who spoken what". Whether it is the conventional modularized approach or the…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Yiling Huang , Weiran Wang , Guanlong Zhao , Hank Liao , Wei Xia , Quan Wang

In this paper, we introduce DiarizationLM, a framework to leverage large language models (LLM) to post-process the outputs from a speaker diarization system. Various goals can be achieved with the proposed framework, such as improving the…

音频与语音处理 · 电气工程与系统科学 2025-01-10 Quan Wang , Yiling Huang , Guanlong Zhao , Evan Clark , Wei Xia , Hank Liao

The LEAP submission for DIHARD-III challenge is described in this paper. The proposed system is composed of a speech bandwidth classifier, and diarization systems fine-tuned for narrowband and wideband speech separately. We use an…

音频与语音处理 · 电气工程与系统科学 2021-06-15 Prachi Singh , Rajat Varma , Venkat Krishnamohan , Srikanth Raj Chetupalli , Sriram Ganapathy

The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as…

声音 · 计算机科学 2022-11-04 You Jin Kim , Hee-Soo Heo , Jee-weon Jung , Youngki Kwon , Bong-Jin Lee , Joon Son Chung

This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target speaker embeddings.…

音频与语音处理 · 电气工程与系统科学 2022-07-14 Weiqing Wang , Qingjian Lin , Ming Li

This paper presents our latest investigation on end-to-end automatic speech recognition (ASR) for overlapped speech. We propose to train an end-to-end system conditioned on speaker embeddings and further improved by transfer learning from…

音频与语音处理 · 电气工程与系统科学 2019-08-14 Pavel Denisov , Ngoc Thang Vu

Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several challenges when adopting EEs for diarisation, from which we…

Majority of speech signals across different scenarios are never available with well-defined audio segments containing only a single speaker. A typical conversation between two speakers consists of segments where their voices overlap,…

音频与语音处理 · 电气工程与系统科学 2022-05-20 Siddharth S. Nijhawan , Homayoon Beigi

Self-supervised models such as WavLM have demonstrated strong performance for neural speaker diarization. However, these models are typically pre-trained on single-channel recordings, limiting their effectiveness in multi-channel scenarios.…

音频与语音处理 · 电气工程与系统科学 2025-10-17 Jiangyu Han , Ruoyu Wang , Yoshiki Masuyama , Marc Delcroix , Johan Rohdin , Jun Du , Lukas Burget

End-to-end speaker diarization approaches have shown exceptional performance over the traditional modular approaches. To further improve the performance of the end-to-end speaker diarization for real speech recordings, recently works have…

声音 · 计算机科学 2022-04-19 Chenyu Yang , Yu Wang

This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target speaker embeddings. By…

音频与语音处理 · 电气工程与系统科学 2023-10-16 Weiqing Wang , Ming Li

Although automatic emotion recognition (AER) has recently drawn significant research interest, most current AER studies use manually segmented utterances, which are usually unavailable for dialogue systems. This paper proposes integrating…

音频与语音处理 · 电气工程与系统科学 2023-08-15 Wen Wu , Chao Zhang , Philip C. Woodland

With the rise in multimedia content over the years, more variety is observed in the recording environments of audio. An audio processing system might benefit when it has a module to identify the acoustic domain at its front-end. In this…

声音 · 计算机科学 2022-08-09 A Kishore Kumar , Shefali Waldekar , Md Sahidullah , Goutam Saha

Overlapped speech detection (OSD) is critical for speech applications in scenario of multi-party conversion. Despite numerous research efforts and progresses, comparing with speech activity detection (VAD), OSD remains an open challenge and…

声音 · 计算机科学 2022-09-27 Ziqing Du , Kai Liu , Xucheng Wan , Huan Zhou

Strong representations of target speakers can help extract important information about speakers and detect corresponding temporal regions in multi-speaker conversations. In this study, we propose a neural architecture that simultaneously…

声音 · 计算机科学 2023-06-07 Chin-Yi Cheng , Hung-Shin Lee , Yu Tsao , Hsin-Min Wang

The DIarization and Speech Processing for LAnguage understanding in Conversational Environments - Medical (DISPLACE-M) challenge introduces a conversational AI benchmark for understanding goal-oriented, real-world medical dialogues. The…

Transformer-based end-to-end neural speaker diarization (EEND) models utilize the multi-head self-attention (SA) mechanism to enable accurate speaker label prediction in overlapped speech regions. In this study, to enhance the training…

音频与语音处理 · 电气工程与系统科学 2023-03-03 Ye-Rin Jeoung , Joon-Young Yang , Jeong-Hwan Choi , Joon-Hyuk Chang

We propose a modular pipeline for the single-channel separation, recognition, and diarization of meeting-style recordings and evaluate it on the Libri-CSS dataset. Using a Continuous Speech Separation (CSS) system with a TF-GridNet…

音频与语音处理 · 电气工程与系统科学 2024-05-07 Thilo von Neumann , Christoph Boeddeker , Tobias Cord-Landwehr , Marc Delcroix , Reinhold Haeb-Umbach

Many approaches can derive information about a single speaker's identity from the speech by learning to recognize consistent characteristics of acoustic parameters. However, it is challenging to determine identity information when there are…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Hyewon Han , Soo-Whan Chung , Hong-Goo Kang