中文
相关论文

相关论文: DiCoW: Diarization-Conditioned Whisper for Target …

200 篇论文

This paper presents Transcribe-to-Diarize, a new approach for neural speaker diarization that uses an end-to-end (E2E) speaker-attributed automatic speech recognition (SA-ASR). The E2E SA-ASR is a joint model that was recently proposed for…

音频与语音处理 · 电气工程与系统科学 2022-01-25 Naoyuki Kanda , Xiong Xiao , Yashesh Gaur , Xiaofei Wang , Zhong Meng , Zhuo Chen , Takuya Yoshioka

We propose a single-channel Deep Cascade Fusion of Diarization and Separation (DCF-DS) framework for back-end automatic speech recognition (ASR), combining neural speaker diarization (NSD) and speech separation (SS). First, we sequentially…

音频与语音处理 · 电气工程与系统科学 2024-12-31 Shu-Tong Niu , Jun Du , Ruo-Yu Wang , Gao-Bin Yang , Tian Gao , Jia Pan , Yu Hu

This paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI. A novel speech-based in-context learning (SICL) approach is proposed for test-time adaptation, which can…

音频与语音处理 · 电气工程与系统科学 2024-03-21 Siyin Wang , Chao-Han Huck Yang , Ji Wu , Chao Zhang

Deep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between all possible speakers, encoding the multiple acoustic…

声音 · 计算机科学 2021-04-26 Chau Luu , Peter Bell , Steve Renals

In this paper, we focus on Whisper, a recent automatic speech recognition model trained with a massive 680k hour labeled speech corpus recorded in diverse conditions. We first show an interesting finding that while Whisper is very robust…

声音 · 计算机科学 2023-10-10 Yuan Gong , Sameer Khurana , Leonid Karlinsky , James Glass

In this paper, we propose a novel end-to-end neural-network-based speaker diarization method. Unlike most existing methods, our proposed method does not have separate modules for extraction and clustering of speaker representations.…

音频与语音处理 · 电气工程与系统科学 2019-09-16 Yusuke Fujita , Naoyuki Kanda , Shota Horiguchi , Kenji Nagamatsu , Shinji Watanabe

We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via…

音频与语音处理 · 电气工程与系统科学 2026-01-13 Mingyue Huo , Yiwen Shao , Yuheng Zhang

Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, prevailing methods augment acoustic…

声音 · 计算机科学 2026-01-28 Xin Zhang , Lin Li , Xiangni Lu , Jianquan Liu , Kong Aik Lee

Deep biasing improves automatic speech recognition (ASR) performance by incorporating contextual phrases. However, most existing methods enhance subwords in a contextual phrase as independent units, potentially compromising contextual…

声音 · 计算机科学 2025-05-30 Zhennan Lin , Kaixun Huang , Wei Ren , Linju Yang , Lei Xie

A crucial part of an accurate and reliable spoken language assessment system is the underlying ASR model. Recently, large-scale pre-trained ASR foundation models such as Whisper have been made available. As the output of these models is…

计算与语言 · 计算机科学 2023-10-11 Rao Ma , Mengjie Qian , Mark J. F. Gales , Kate M. Knill

The performance of speech enhancement algorithms in a multi-speaker scenario depends on correctly identifying the target speaker to be enhanced. Auditory attention decoding (AAD) methods allow to identify the target speaker which the…

声音 · 计算机科学 2020-05-12 Ali Aroudi , Marc Delcroix , Tomohiro Nakatani , Keisuke Kinoshita , Shoko Araki , Simon Doclo

We propose Speaker-Conditioned Serialized Output Training (SC-SOT), an enhanced SOT-based training for E2E multi-talker ASR. We first probe how SOT handles overlapped speech, and we found the decoder performs implicit speaker separation. We…

声音 · 计算机科学 2025-06-17 Yuta Hirano , Sakriani Sakti

Whispering is a distinct form of speech known for its soft, breathy, and hushed characteristics, often used for private communication. The acoustic characteristics of whispered speech differ substantially from normally phonated speech and…

音频与语音处理 · 电气工程与系统科学 2024-02-08 Zhaofeng Lin , Tanvina Patel , Odette Scharenborg

Recent success in speech representation learning enables a new way to leverage unlabeled data to train speech recognition model. In speech representation learning, a large amount of unlabeled data is used in a self-supervised manner to…

音频与语音处理 · 电气工程与系统科学 2020-12-15 Shaoshi Ling , Yuzong Liu

Diffusion-based large language models (DLLMs) have recently attracted growing interest as an alternative to autoregressive decoders. In this work, we present an empirical study on using the diffusion-based large language model LLaDA for…

音频与语音处理 · 电气工程与系统科学 2026-03-02 Mengqi Wang , Zhan Liu , Zengrui Jin , Guangzhi Sun , Chao Zhang , Philip C. Woodland

In this work, we propose an error correction framework, named DiaCorrect, to refine the output of a diarization system in a simple yet effective way. This method is inspired by error correction techniques in automatic speech recognition.…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Jiangyu Han , Federico Landini , Johan Rohdin , Mireia Diez , Lukas Burget , Yuhang Cao , Heng Lu , Jan Cernocky

This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target speaker embeddings.…

音频与语音处理 · 电气工程与系统科学 2022-07-14 Weiqing Wang , Qingjian Lin , Ming Li

In the task of speaker diarization, the number of small-scale meetings accounts for a large proportion. When microphone arrays are employed as a recording device, its spatial information is usually ignored by most researchers. In this…

声音 · 计算机科学 2022-10-27 Yuxuan Du , Ruohua Zhou

This paper presents a novel framework for joint speaker diarization (SD) and automatic speech recognition (ASR), named SLIDAR (sliding-window diarization-augmented recognition). SLIDAR can process arbitrary length inputs and can handle any…

音频与语音处理 · 电气工程与系统科学 2023-10-04 Samuele Cornell , Jee-weon Jung , Shinji Watanabe , Stefano Squartini

Speaker identification in noisy audio recordings, specifically those from collaborative learning environments, can be extremely challenging. There is a need to identify individual students talking in small groups from other students talking…

音频与语音处理 · 电气工程与系统科学 2022-07-05 Antonio Gomez