English
Related papers

Related papers: Deep MOS Predictor for Synthetic Speech Using Clus…

200 papers

Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we…

Neural Text-to-speech (TTS) synthesis is a powerful technology that can generate speech using neural networks. One of the most remarkable features of TTS synthesis is its capability to produce speech in the voice of different speakers. This…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-19 Vinotha R , Hepsiba D , L. D. Vijay Anand , Deepak John Reji

We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-04 Dyah A. M. G. Wisnu , Ryandhimas E. Zezario , Stefano Rini , Hsin-Min Wang , Yu Tsao

We design an online end-to-end speech recognition system based on Time-Depth Separable (TDS) convolutions and Connectionist Temporal Classification (CTC). We improve the core TDS architecture in order to limit the future context and hence…

Traditionally, research in automated speech recognition has focused on local-first encoding of audio representations to predict the spoken phonemes in an utterance. Unfortunately, approaches relying on such hyper-local information tend to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-19 David M. Chan , Shalini Ghosh , Debmalya Chakrabarty , Björn Hoffmeister

This paper proposes a novel approach that uses deep neural networks for classifying imagined speech, significantly increasing the classification accuracy. The proposed approach employs only the EEG channels over specific areas of the brain…

Neurons and Cognition · Quantitative Biology 2020-03-24 Jerrin Thomas Panachakel , A. G. Ramakrishnan , A. G. Ramakrishnan

We propose a multi-stage framework for universal speech enhancement, designed for the Interspeech 2025 URGENT Challenge. Our system first employs a Sparse Compression Network to robustly separate sources and extract an initial clean speech…

Sound · Computer Science 2025-06-03 Nabarun Goswami , Tatsuya Harada

This paper discusses the task of face-based speech synthesis, a kind of personalized speech synthesis where the synthesized voices are constrained to perceptually match with a reference face image. Due to the lack of TTS-quality…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-07 Yao Shi , Yunfei Xu , Hongbin Suo , Yulong Wan , Haifeng Liu

The performance of most speaker diarization systems with x-vector embeddings is both vulnerable to noisy environments and lacks domain robustness. Earlier work on speaker diarization using generative adversarial network (GAN) with an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-21 Monisankha Pal , Manoj Kumar , Raghuveer Peri , Tae Jin Park , So Hyun Kim , Catherine Lord , Somer Bishop , Shrikanth Narayanan

The Mean Opinion Score (MOS) serves as the standard metric for speech quality assessment, yet biases in human annotations remain underexplored. We conduct the first systematic analysis of gender bias in MOS, revealing that male listeners…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Wenze Ren , Yi-Cheng Lin , Wen-Chin Huang , Erica Cooper , Ryandhimas E. Zezario , Hsin-Min Wang , Hung-yi Lee , Yu Tsao

Deep Neural Networks (DNN) have been successful in en- hancing noisy speech signals. Enhancement is achieved by learning a nonlinear mapping function from the features of the corrupted speech signal to that of the reference clean speech…

Machine Learning · Computer Science 2016-06-16 Zhenzhou Wu , Sunil Sivadas , Yong Kiam Tan , Ma Bin , Rick Siow Mong Goh

This paper aims to enhance low-resource TTS by reducing training data requirements using compact speech representations. A Multi-Stage Multi-Codebook (MSMC) VQ-GAN is trained to learn the representation, MSMCR, and decode it to waveforms.…

Sound · Computer Science 2022-10-28 Haohan Guo , Fenglong Xie , Xixin Wu , Hui Lu , Helen Meng

Recent research on speech enhancement (SE) has seen the emergence of deep-learning-based methods. It is still a challenging task to determine the effective ways to increase the generalizability of SE under diverse test conditions. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-01 Ryandhimas E. Zezario , Chiou-Shann Fuh , Hsin-Min Wang , Yu Tsao

Neural speech synthesis models have recently demonstrated the ability to synthesize high quality speech for text-to-speech and compression applications. These new models often require powerful GPUs to achieve real-time operation, so being…

Audio and Speech Processing · Electrical Eng. & Systems 2019-02-20 Jean-Marc Valin , Jan Skoglund

The mean opinion score (MOS) is a standard metric for assessing speech quality, but its singular focus fails to identify specific distortions when low scores are observed. The NISQA dataset addresses this limitation by providing ratings…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-06 Fredrik Cumlin , Xinyu Liang , Victor Ungureanu , Chandan K. A. Reddy , Christian Schüldt , Saikat Chatterjee

Computational modeling of naturalistic conversations in clinical applications has seen growing interest in the past decade. An important use-case involves child-adult interactions within the autism diagnosis and intervention domain. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-30 Nithin Rao Koluguri , Manoj Kumar , So Hyun Kim , Catherine Lord , Shrikanth Narayanan

This paper explores sequential modelling of polyphonic music with deep neural networks. While recent breakthroughs have focussed on network architecture, we demonstrate that the representation of the sequence can make an equally significant…

Sound · Computer Science 2021-08-11 Omar Peracha

In recent years, Text-To-Speech (TTS) has been used as a data augmentation technique for speech recognition to help complement inadequacies in the training data. Correspondingly, we investigate the use of a multi-speaker TTS system to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-25 Yiling Huang , Yutian Chen , Jason Pelecanos , Quan Wang

Most of the current deep learning-based approaches for speech enhancement only operate in the spectrogram or waveform domain. Although a cross-domain transformer combining waveform- and spectrogram-domain inputs has been proposed, its…

Sound · Computer Science 2023-10-31 Jialu Li , Junhui Li , Pu Wang , Youshan Zhang

Automatic classification of disordered speech can provide an objective tool for identifying the presence and severity of speech impairment. Classification approaches can also help identify hard-to-recognize speech samples to teach ASR…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-09 Subhashini Venugopalan , Joel Shor , Manoj Plakal , Jimmy Tobin , Katrin Tomanek , Jordan R. Green , Michael P. Brenner
‹ Prev 1 8 9 10 Next ›