English
Related papers

Related papers: Deep Speech 2: End-to-End Speech Recognition in En…

200 papers

Building cross-lingual voice conversion (VC) systems for multiple speakers and multiple languages has been a challenging task for a long time. This paper describes a parallel non-autoregressive network to achieve bilingual and code-switched…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-23 Yaogen Yang , Haozhe Zhang , Xiaoyi Qin , Shanshan Liang , Huahua Cui , Mingyang Xu , Ming Li

Transformers are powerful neural architectures that allow integrating different modalities using attention mechanisms. In this paper, we leverage the neural transformer architectures for multi-channel speech recognition systems, where the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Feng-Ju Chang , Martin Radfar , Athanasios Mouchtaris , Brian King , Siegfried Kunzmann

The attention-based deep contextual biasing method has been demonstrated to effectively improve the recognition performance of end-to-end automatic speech recognition (ASR) systems on given contextual phrases. However, unlike shallow fusion…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-10 Kaixun Huang , Ao Zhang , Binbin Zhang , Tianyi Xu , Xingchen Song , Lei Xie

A novel interpretable end-to-end learning scheme for language identification is proposed. It is in line with the classical GMM i-vector methods both theoretically and practically. In the end-to-end pipeline, a general encoding layer is…

Audio and Speech Processing · Electrical Eng. & Systems 2018-04-03 Weicheng Cai , Zexin Cai , Wenbo Liu , Xiaoqi Wang , Ming Li

In this paper, we investigate how the output representation of an end-to-end neural network affects multilingual automatic speech recognition (ASR). We study different representations including character-level, byte-level, byte pair…

Computation and Language · Computer Science 2022-05-03 Liuhui Deng , Roger Hsiao , Arnab Ghoshal

Automatic speech recognition (ASR) systems degrade significantly under noisy conditions. Recently, speech enhancement (SE) is introduced as front-end to reduce noise for ASR, but it also suppresses some important speech information, i.e.,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Yuchen Hu , Nana Hou , Chen Chen , Eng Siong Chng

Recently, there has been an increasing interest in neural speech synthesis. While the deep neural network achieves the state-of-the-art result in text-to-speech (TTS) tasks, how to generate a more emotional and more expressive speech is…

Computation and Language · Computer Science 2021-06-24 Chenye Cui , Yi Ren , Jinglin Liu , Feiyang Chen , Rongjie Huang , Ming Lei , Zhou Zhao

Current front-ends for robust automatic speech recognition(ASR) include masking- and mapping-based deep learning approaches to speech enhancement. A recently proposed deep learning approach toa prioriSNR estimation, called DeepXi, was able…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-29 Aaron Nicolson , Kuldip K. Paliwal

The field of speech processing has undergone a transformative shift with the advent of deep learning. The use of multiple processing layers has enabled the creation of models capable of extracting intricate features from speech data. This…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-31 Ambuj Mehrish , Navonil Majumder , Rishabh Bhardwaj , Rada Mihalcea , Soujanya Poria

Using end-to-end models for speech translation (ST) has increasingly been the focus of the ST community. These models condense the previously cascaded systems by directly converting sound waves into translated text. However, cascaded models…

Computation and Language · Computer Science 2021-01-25 Orion Weller , Matthias Sperber , Christian Gollan , Joris Kluivers

In the last decade of automatic speech recognition (ASR) research, the introduction of deep learning brought considerable reductions in word error rate of more than 50% relative, compared to modeling without deep learning. In the wake of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-07 Rohit Prabhavalkar , Takaaki Hori , Tara N. Sainath , Ralf Schlüter , Shinji Watanabe

The goal of this work is to synchronise audio and video of a talking face using deep neural network models. Existing works have trained networks on proxy tasks such as cross-modal similarity learning, and then computed similarities between…

Computer Vision and Pattern Recognition · Computer Science 2021-03-22 You Jin Kim , Hee Soo Heo , Soo-Whan Chung , Bong-Jin Lee

Recent diarization technologies can be categorized into two approaches, i.e., clustering and end-to-end neural approaches, which have different pros and cons. The clustering-based approaches assign speaker labels to speech regions by…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-08 Keisuke Kinoshita , Marc Delcroix , Naohiro Tawara

The Mandarin Chinese language is known to be strongly influenced by a rich set of regional accents, while Mandarin speech with each accent is quite low resource. Hence, an important task in Mandarin speech recognition is to appropriately…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-17 Xurong Xie , Xiang Sui , Xunying Liu , Lan Wang

To scale neural speech synthesis to various real-world languages, we present a multilingual end-to-end framework that maps byte inputs to spectrograms, thus allowing arbitrary input scripts. Besides strong results on 40+ languages, the…

Computation and Language · Computer Science 2021-07-12 Mutian He , Jingzhou Yang , Lei He , Frank K. Soong

Techniques for multi-lingual and cross-lingual speech recognition can help in low resource scenarios, to bootstrap systems and enable analysis of new languages and domains. End-to-end approaches, in particular sequence-based techniques, are…

Computation and Language · Computer Science 2018-03-08 Siddharth Dalmia , Ramon Sanabria , Florian Metze , Alan W. Black

In this paper, we explore the encoding/pooling layer and loss function in the end-to-end speaker and language recognition system. First, a unified and interpretable end-to-end system for both speaker and language recognition is developed.…

Audio and Speech Processing · Electrical Eng. & Systems 2018-04-17 Weicheng Cai , Jinkun Chen , Ming Li

End-to-end automatic speech recognition (ASR) models, including both attention-based models and the recurrent neural network transducer (RNN-T), have shown superior performance compared to conventional systems. However, previous studies…

Despite the increasing research interest in end-to-end learning systems for speech emotion recognition, conventional systems either suffer from the overfitting due in part to the limited training data, or do not explicitly consider the…

Computation and Language · Computer Science 2019-04-01 Zixing Zhang , Bingwen Wu , Bjoern Schuller

While efficient architectures and a plethora of augmentations for end-to-end image classification tasks have been suggested and heavily investigated, state-of-the-art techniques for audio classifications still rely on numerous…

Sound · Computer Science 2022-07-06 Avi Gazneli , Gadi Zimerman , Tal Ridnik , Gilad Sharir , Asaf Noy