中文
相关论文

相关论文: Mixtures of Deep Neural Experts for Automated Spee…

200 篇论文

Second language proficiency (L2) in English is usually perceptually evaluated by English teachers or expert evaluators, with the inherent intra- and inter-rater variability. This paper explores deep learning techniques for comprehensive L2…

计算与语言 · 计算机科学 2025-05-06 Armita Mohammadi , Alessandro Lameiras Koerich , Laureano Moro-Velazquez , Patrick Cardinal

Automatic Speech Scoring (ASS) is the computer-assisted evaluation of a candidate's speaking proficiency in a language. ASS systems face many challenges like open grammar, variable pronunciations, and unstructured or semi-structured…

音频与语音处理 · 电气工程与系统科学 2021-09-07 Yaman Kumar Singla , Avykat Gupta , Shaurya Bagga , Changyou Chen , Balaji Krishnamurthy , Rajiv Ratn Shah

We introduce a new automatic evaluation method for speaker similarity assessment, that is consistent with human perceptual scores. Modern neural text-to-speech models require a vast amount of clean training data, which is why many solutions…

声音 · 计算机科学 2022-07-04 Deja Kamil , Sanchez Ariadna , Roth Julian , Cotescu Marius

This paper presents a novel deep learning architecture for acoustic model in the context of Automatic Speech Recognition (ASR), termed as MixNet. Besides the conventional layers, such as fully connected layers in DNN-HMM and memory cells in…

音频与语音处理 · 电气工程与系统科学 2021-12-03 Vishwanath Pratap Singh , Shakti P. Rath , Abhishek Pandey

Automatic speech quality assessment plays a crucial role in the development of speech synthesis systems, but existing models exhibit significant performance variations across different granularity levels of prediction tasks. This paper…

声音 · 计算机科学 2025-07-09 Xintong Hu , Yixuan Chen , Rui Yang , Wenxiang Guo , Changhao Pan

Mixtures of Experts combine the outputs of several "expert" networks, each of which specializes in a different part of the input space. This is achieved by training a "gating" network that maps each input to a distribution over the experts.…

机器学习 · 计算机科学 2014-03-11 David Eigen , Marc'Aurelio Ranzato , Ilya Sutskever

This paper presents a study on discriminative artificial neural network classifiers in the context of open-set speaker identification. Both 2-class and multi-class architectures are tested against the conventional Gaussian mixture model…

机器学习 · 计算机科学 2019-04-03 Stefano Imoscopi , Volodya Grancharov , Sigurdur Sverrisson , Erlendur Karlsson , Harald Pobloth

While neural networks have been employed to handle several different text-to-speech tasks, ours is the first system to use neural networks throughout, for both linguistic and acoustic processing. We divide the text-to-speech task into three…

神经与进化计算 · 计算机科学 2016-11-17 Orhan Karaali , Gerald Corrigan , Noel Massey , Corey Miller , Otto Schnurr , Andrew Mackie

In this paper, we investigate a deep learning approach for speech denoising through an efficient ensemble of specialist neural networks. By splitting up the speech denoising task into non-overlapping subproblems and introducing a…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Aswin Sivaraman , Minje Kim

Speech deepfakes pose a significant threat to personal security and content authenticity. Several detectors have been proposed in the literature, and one of the primary challenges these systems have to face is the generalization over unseen…

声音 · 计算机科学 2024-09-25 Viola Negroni , Davide Salvi , Alessandro Ilic Mezza , Paolo Bestagini , Stefano Tubaro

Multi-lingual speech recognition aims to distinguish linguistic expressions in different languages and integrate acoustic processing simultaneously. In contrast, current multi-lingual speech recognition research follows a language-aware…

音频与语音处理 · 电气工程与系统科学 2023-02-28 Yoohwan Kwon , Soo-Whan Chung

Second pass rescoring is a critical component of competitive automatic speech recognition (ASR) systems. Large language models have demonstrated their ability in using pre-trained information for better rescoring of ASR hypothesis.…

音频与语音处理 · 电气工程与系统科学 2023-10-11 Prashanth Gurunath Shivakumar , Jari Kolehmainen , Yile Gu , Ankur Gandhe , Ariya Rastrow , Ivan Bulyko

In this study we present a Deep Mixture of Experts (DMoE) neural-network architecture for single microphone speech enhancement. By contrast to most speech enhancement algorithms that overlook the speech variability mainly caused by phoneme…

声音 · 计算机科学 2017-03-29 Shlomo E. Chazan , Jacob Goldberger , Sharon Gannot

Natural and artificial audition can in principle acquire different solutions to a given problem. The constraints of the task, however, can nudge the cognitive science and engineering of audition to qualitatively converge, suggesting that a…

声音 · 计算机科学 2023-04-20 Federico Adolfi , Jeffrey S. Bowers , David Poeppel

Multilingual speaker verification introduces the challenge of verifying a speaker in multiple languages. Existing systems were built using i-vector/x-vector approaches along with Bi-LSTMs, which were trained to discriminate speakers,…

声音 · 计算机科学 2024-08-09 Aravinda Reddy PN , Raghavendra Ramachandra , K. Sreenivasa Rao , Pabitra Mitra

Speech evaluation measures a learners oral proficiency using automatic models. Corpora for training such models often pose sparsity challenges given that there often is limited scored data from teachers, in addition to the score…

人工智能 · 计算机科学 2024-09-24 Huayun Zhang , Jeremy H. M. Wong , Geyu Lin , Nancy F. Chen

Self-supervised learning models have revolutionized the field of speech processing. However, the process of fine-tuning these models on downstream tasks requires substantial computational resources, particularly when dealing with multiple…

计算与语言 · 计算机科学 2024-06-24 Varsha Suresh , Salah Aït-Mokhtar , Caroline Brun , Ioan Calapodescu

Multi-source localization is an important and challenging technique for multi-talker conversation analysis. This paper proposes a novel supervised learning method using deep neural networks to estimate the direction of arrival (DOA) of all…

音频与语音处理 · 电气工程与系统科学 2021-11-30 Aswin Shanmugam Subramanian , Chao Weng , Shinji Watanabe , Meng Yu , Dong Yu

Level assessment for foreign language students is necessary for putting them in the right level group, furthermore, interviewing students is a very time-consuming task, so we propose to automate the evaluation of speaker fluency level by…

机器学习 · 统计学 2018-09-03 Alan Preciado-Grijalva , Ramon F. Brena

Most state-of-the-art Deep Learning systems for speaker verification are based on speaker embedding extractors. These architectures are commonly composed of a feature extractor front-end together with a pooling layer to encode…

音频与语音处理 · 电气工程与系统科学 2021-01-12 Miquel India , Pooyan Safari , Javier Hernando
‹ 上一页 1 2 3 10 下一页 ›