中文
相关论文

相关论文: Serialized Output Training by Learned Dominance

200 篇论文

Speech separation has been well developed, with the very successful permutation invariant training (PIT) approach, although the frequent label assignment switching happening during PIT training remains to be a problem when better…

声音 · 计算机科学 2021-08-24 Sung-Feng Huang , Shun-Po Chuang , Da-Rong Liu , Yi-Chen Chen , Gene-Ping Yang , Hung-yi Lee

Streaming recognition of multi-talker conversations has so far been evaluated only for 2-speaker single-turn sessions. In this paper, we investigate it for multi-turn meetings containing multiple speakers using the Streaming Unmixing and…

音频与语音处理 · 电气工程与系统科学 2022-01-25 Desh Raj , Liang Lu , Zhuo Chen , Yashesh Gaur , Jinyu Li

Spoken language understanding (SLU) requires a model to analyze input acoustic signal to understand its linguistic content and make predictions. To boost the models' performance, various pre-training methods have been proposed to learn rich…

计算与语言 · 计算机科学 2021-03-16 Yu-An Chung , Chenguang Zhu , Michael Zeng

Permutation Invariant Training (PIT) has long been a stepping stone method for training speech separation model in handling the label ambiguity problem. With PIT selecting the minimum cost label assignments dynamically, very few studies…

声音 · 计算机科学 2019-10-29 Gene-Ping Yang , Szu-Lin Wu , Yao-Wen Mao , Hung-yi Lee , Lin-shan Lee

Self-supervised learning (SSL) on large-scale datasets like AudioSet has become the dominant paradigm for audio representation learning. While the continuous influx of new, unlabeled audio presents an opportunity to enrich these static…

声音 · 计算机科学 2026-01-26 Yizhou Zhang , Yuan Gao , Wangjin Zhou , Zicheng Yuan , Keisuke Imoto , Tatsuya Kawahara

Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…

音频与语音处理 · 电气工程与系统科学 2024-09-10 Hongyu Wang , Hui Li , Bo Li

Tonal low-resource languages are widely spoken yet remain underserved by modern speech technology. A key challenge is learning representations that are robust to nuisance variation such as gender while remaining tone-aware for different…

计算与语言 · 计算机科学 2026-01-15 Tianyi Xu , Xuan Ouyang , Binwei Yao , Shoua Xiong , Sara Misurelli , Maichou Lor , Junjie Hu

While Transformers have achieved remarkable success in LLMs through superior scalability, their application in industrial-scale ranking models remains nascent, hindered by the challenges of high feature sparsity and low label density. In…

信息检索 · 计算机科学 2026-03-05 Chunqi Wang , Bingchao Wu , Taotian Pang , Jiahao Wang , Jie Yang , Jia Liu , Hao Zhang , Hai Zhu , Lei Shen , Shizhun Wang , Bing Wang , Xiaoyi Zeng

An important problem in machine auditory perception is to recognize and detect sound events. In this paper, we propose a sequential self-teaching approach to learning sounds. Our main proposition is that it is harder to learn sounds in…

声音 · 计算机科学 2020-07-02 Anurag Kumar , Vamsi Krishna Ithapu

Self-training (ST) and self-supervised learning (SSL) methods have demonstrated strong improvements in automatic speech recognition (ASR). In spite of these advances, to the best of our knowledge, there is no analysis of how the composition…

机器学习 · 计算机科学 2023-03-03 Dan Berrebbi , Ronan Collobert , Navdeep Jaitly , Tatiana Likhomanenko

Unsupervised object-centric learning aims to decompose scenes into interpretable object entities, termed slots. Slot-based auto-encoders stand out as a prominent method for this task. Within them, crucial aspects include guiding the encoder…

计算机视觉与模式识别 · 计算机科学 2024-04-08 Ioannis Kakogeorgiou , Spyros Gidaris , Konstantinos Karantzalos , Nikos Komodakis

We propose a novel adversarial multi-task learning scheme, aiming at actively curtailing the inter-talker feature variability while maximizing its senone discriminability so as to enhance the performance of a deep neural network (DNN) based…

音频与语音处理 · 电气工程与系统科学 2019-05-01 Zhong Meng , Jinyu Li , Zhuo Chen , Yong Zhao , Vadim Mazalov , Yifan Gong , Biing-Hwang , Juang

This paper proposes a serialized multi-layer multi-head attention for neural speaker embedding in text-independent speaker verification. In prior works, frame-level features from one layer are aggregated to form an utterance-level…

声音 · 计算机科学 2021-07-15 Hongning Zhu , Kong Aik Lee , Haizhou Li

Multi-talker conversational speech processing has drawn many interests for various applications such as meeting transcription. Speech separation is often required to handle overlapped speech that is commonly observed in conversation.…

音频与语音处理 · 电气工程与系统科学 2021-11-18 Wangyou Zhang , Zhuo Chen , Naoyuki Kanda , Shujie Liu , Jinyu Li , Sefik Emre Eskimez , Takuya Yoshioka , Xiong Xiao , Zhong Meng , Yanmin Qian , Furu Wei

In this work, we propose a novel method for modeling numerous speakers, which enables expressing the overall characteristics of speakers in detail like a trained multi-speaker model without additional training on the target speaker's…

声音 · 计算机科学 2024-06-03 Jungil Kong , Junmo Lee , Jeongmin Kim , Beomjeong Kim , Jihoon Park , Dohee Kong , Changheon Lee , Sangjin Kim

Multi-talker speech recognition (MTASR) faces unique challenges in disentangling and transcribing overlapping speech. To address these challenges, this paper investigates the role of Connectionist Temporal Classification (CTC) in speaker…

音频与语音处理 · 电气工程与系统科学 2025-01-06 Jiawen Kang , Lingwei Meng , Mingyu Cui , Yuejiao Wang , Xixin Wu , Xunying Liu , Helen Meng

Recent studies show that training deep neural networks (DNNs) with Lipschitz constraints are able to enhance adversarial robustness and other model properties such as stability. In this paper, we propose a layer-wise orthogonal training…

机器学习 · 计算机科学 2023-03-28 Xiaojun Xu , Linyi Li , Bo Li

The Streaming Unmixing and Recognition Transducer (SURT) model was proposed recently as an end-to-end approach for continuous, streaming, multi-talker speech recognition (ASR). Despite impressive results on multi-turn meetings, SURT has…

音频与语音处理 · 电气工程与系统科学 2023-09-20 Desh Raj , Daniel Povey , Sanjeev Khudanpur

Permutation invariant training (PIT) is a widely used training criterion for neural network-based source separation, used for both utterance-level separation with utterance-level PIT (uPIT) and separation of long recordings with the…

音频与语音处理 · 电气工程与系统科学 2021-08-02 Thilo von Neumann , Christoph Boeddeker , Keisuke Kinoshita , Marc Delcroix , Reinhold Haeb-Umbach

Attention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because the decoder predicts text tokens (such as characters or words) in an autoregressive manner, it is difficult for an AED…

计算与语言 · 计算机科学 2021-08-31 Ye Bai , Jiangyan Yi , Jianhua Tao , Zhengkun Tian , Zhengqi Wen , Shuai Zhang