中文
相关论文

相关论文: Location-based training for multi-channel talker-i…

200 篇论文

The performance of automatic speech recognition (ASR) systems severely degrades when multi-talker speech overlap occurs. In meeting environments, speech separation is typically performed to improve the robustness of ASR systems. Recently,…

音频与语音处理 · 电气工程与系统科学 2023-01-18 Hassan Taherian , DeLiang Wang

We study permutation invariant training (PIT), which targets at the permutation ambiguity problem for speaker independent source separation models. We extend two state-of-the-art PIT strategies. First, we look at the two-stage speaker…

声音 · 计算机科学 2021-04-06 Xiaoyu Liu , Jordi Pons

Deep learning has shown a great potential for speech separation, especially for speech and non-speech separation. However, it encounters permutation problem for multi-speaker separation where both target and interference are speech.…

声音 · 计算机科学 2021-03-29 Hao Li , Xueliang Zhang , Guanglai Gao

The goal of speech separation is to extract multiple speech sources from a single microphone recording. Recently, with the advancement of deep learning and availability of large datasets, speech separation has been formulated as a…

音频与语音处理 · 电气工程与系统科学 2021-11-17 Midia Yousefi , John H. L. Hansen

Training speech separation models in the supervised setting raises a permutation problem: finding the best assignation between the model predictions and the ground truth separated signals. This inherently ambiguous task is customarily…

声音 · 计算机科学 2024-11-28 David Perera , François Derrida , Théo Mariotte , Gaël Richard , Slim Essid

Single channel speech separation has experienced great progress in the last few years. However, training neural speech separation for a large number of speakers (e.g., more than 10 speakers) is out of reach for the current methods, which…

声音 · 计算机科学 2021-11-09 Shaked Dovrat , Eliya Nachmani , Lior Wolf

In this paper we propose the utterance-level Permutation Invariant Training (uPIT) technique. uPIT is a practically applicable, end-to-end, deep learning based solution for speaker independent multi-talker speech separation. Specifically,…

声音 · 计算机科学 2018-12-06 Morten Kolbæk , Dong Yu , Zheng-Hua Tan , Jesper Jensen

We propose a novel deep learning model, which supports permutation invariant training (PIT), for speaker independent multi-talker speech separation, commonly known as the cocktail-party problem. Different from most of the prior arts that…

计算与语言 · 计算机科学 2018-12-06 Dong Yu , Morten Kolbæk , Zheng-Hua Tan , Jesper Jensen

Single-microphone, speaker-independent speech separation is normally performed through two steps: (i) separating the specific speech sources, and (ii) determining the best output-label assignment to find the separation error. The second…

音频与语音处理 · 电气工程与系统科学 2019-08-07 Midia Yousefi , Soheil Khorram , John H. L. Hansen

Permutation Invariant Training (PIT) has long been a stepping stone method for training speech separation model in handling the label ambiguity problem. With PIT selecting the minimum cost label assignments dynamically, very few studies…

声音 · 计算机科学 2019-10-29 Gene-Ping Yang , Szu-Lin Wu , Yao-Wen Mao , Hung-yi Lee , Lin-shan Lee

Utterance-level permutation invariant training (uPIT) has achieved promising progress on single-channel multi-talker speech separation task. Long short-term memory (LSTM) and bidirectional LSTM (BLSTM) are widely used as the separation…

声音 · 计算机科学 2019-12-30 Lu Huang , Gaofeng Cheng , Pengyuan Zhang , Yi Yang , Shumin Xu , Jiasong Sun

Permutation invariant training (PIT) is a widely used training criterion for neural network-based source separation, used for both utterance-level separation with utterance-level PIT (uPIT) and separation of long recordings with the…

音频与语音处理 · 电气工程与系统科学 2021-08-02 Thilo von Neumann , Christoph Boeddeker , Keisuke Kinoshita , Marc Delcroix , Reinhold Haeb-Umbach

Multi-talker conversational speech processing has drawn many interests for various applications such as meeting transcription. Speech separation is often required to handle overlapped speech that is commonly observed in conversation.…

音频与语音处理 · 电气工程与系统科学 2021-11-18 Wangyou Zhang , Zhuo Chen , Naoyuki Kanda , Shujie Liu , Jinyu Li , Sefik Emre Eskimez , Takuya Yoshioka , Xiong Xiao , Zhong Meng , Yanmin Qian , Furu Wei

Speech separation has been studied in time domain because of lower latency and higher performance compared to time-frequency domain. The masking-based method has been mostly used in time domain, and the other common method (mapping-based)…

声音 · 计算机科学 2022-03-22 Chenyang Gao , Yue Gu , Ivan Marsic

Deep clustering (DC) and utterance-level permutation invariant training (uPIT) have been demonstrated promising for speaker-independent speech separation. DC is usually formulated as two-step processes: embedding learning and embedding…

声音 · 计算机科学 2019-07-24 Cunhang Fan , Bin Liu , Jianhua Tao , Jiangyan Yi , Zhengqi Wen

Universal sound separation consists of separating mixes with arbitrary sounds of different types, and permutation invariant training (PIT) is used to train source agnostic models that do so. In this work, we complement PIT with adversarial…

声音 · 计算机科学 2023-03-07 Emilian Postolache , Jordi Pons , Santiago Pascual , Joan Serrà

In neural network-based monaural speech separation techniques, it has been recently common to evaluate the loss using the permutation invariant training (PIT) loss. However, the ordinary PIT requires to try all $N!$ permutations between $N$…

声音 · 计算机科学 2021-05-18 Hideyuki Tachibana

In this paper, we propose a novel technique for direct recognition of multiple speech streams given the single channel of mixed speech, without first separating them. Our technique is based on permutation invariant training (PIT) for…

声音 · 计算机科学 2018-12-06 Dong Yu , Xuankai Chang , Yanmin Qian

In supervised speech separation, permutation invariant training (PIT) is widely used to handle label ambiguity by selecting the best permutation to update the model. Despite its success, previous studies showed that PIT is plagued by…

声音 · 计算机科学 2023-11-22 Chenyang Gao , Yue Gu , Ivan Marsic

In this paper, we propose a novel strategy for text-independent speaker identification system: Multi-Label Training (MLT). Instead of the commonly used one-to-one correspondence between the speech and the speaker label, we divide all the…

音频与语音处理 · 电气工程与系统科学 2024-08-19 Yuqi Xue
‹ 上一页 1 2 3 10 下一页 ›