English
Related papers

Related papers: An Objective Evaluation Framework for Pathological…

200 papers

Speech synthesis has come a long way as current text-to-speech (TTS) models can now generate natural human-sounding speech. However, most of the TTS research focuses on using adult speech data and there has been very limited work done on…

Sound · Computer Science 2022-04-05 Rishabh Jain , Mariam Yiwere , Dan Bigioi , Peter Corcoran , Horia Cucu

The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our study applies a modified Transformer in a speech enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-04 Szu-Wei Fu , Chien-Feng Liao , Tsun-An Hsieh , Kuo-Hsuan Hung , Syu-Siang Wang , Cheng Yu , Heng-Cheng Kuo , Ryandhimas E. Zezario , You-Jin Li , Shang-Yi Chuang , Yen-Ju Lu , Yu Tsao

Automatic recognition of disordered speech remains a highly challenging task to date. Sources of variability commonly found in normal speech including accent, age or gender, when further compounded with the underlying causes of speech…

Sound · Computer Science 2022-01-20 Mengzhe Geng , Shansong Liu , Jianwei Yu , Xurong Xie , Shoukang Hu , Zi Ye , Zengrui Jin , Xunying Liu , Helen Meng

The rise of AI-driven generative models has enabled the creation of highly realistic speech deepfakes - synthetic audio signals that can imitate target speakers' voices - raising critical security concerns. Existing methods for detecting…

Sound · Computer Science 2025-03-25 Emma Coletta , Davide Salvi , Viola Negroni , Daniele Ugo Leonzio , Paolo Bestagini

Accent plays a significant role in speech communication, influencing one's capability to understand as well as conveying a person's identity. This paper introduces a novel and efficient framework for accented Text-to-Speech (TTS) synthesis…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-01 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

Recent advances in deep learning and computer vision have made the synthesis and counterfeiting of multimedia content more accessible than ever, leading to possible threats and dangers from malicious users. In the audio field, we are…

Sound · Computer Science 2023-07-31 Daniele Mari , Davide Salvi , Paolo Bestagini , Simone Milani

Dysarthric speech exhibits abnormal prosody and significant speaker variability, presenting persistent challenges for automatic speech recognition (ASR). While text-to-speech (TTS)-based data augmentation has shown potential, existing…

Sound · Computer Science 2026-03-03 Minghui Wu , Xueling Liu , Jiahuan Fan , Haitao Tang , Yanyong Zhang , Yue Zhang

This work defines a new framework for performance evaluation of polyphonic sound event detection (SED) systems, which overcomes the limitations of the conventional collar-based event decisions, event F-scores and event error rates. The…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-17 Cagdas Bilen , Giacomo Ferroni , Francesco Tuveri , Juan Azcarreta , Sacha Krstulovic

Speech resynthesis is a generic task for which we want to synthesize audio with another audio as input, which finds applications for media monitors and journalists.Among different tasks addressed by speech resynthesis, voice conversion…

Sound · Computer Science 2024-08-07 Thibault Gaudier , Marie Tahon , Anthony Larcher , Yannick Estève

With the huge technological advances introduced by deep learning in audio & speech processing, many novel synthetic speech techniques achieved incredible realistic results. As these methods generate realistic fake human voices, they can be…

Traditional voice conversion(VC) has been focused on speaker identity conversion for speech with a neutral expression. We note that emotional expression plays an essential role in daily communication, and the emotional style of speech can…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-22 Zongyang Du , Berrak Sisman , Kun Zhou , Haizhou Li

State-of-the-art statistical parametric speech synthesis (SPSS) generally uses a vocoder to represent speech signals and parameterize them into features for subsequent modeling. Magnitude spectrum has been a dominant feature over the years.…

Sound · Computer Science 2015-10-08 Bo Fan , Siu Wa Lee , Xiaohai Tian , Lei Xie , Minghui Dong

Speaker verification systems are vulnerable to spoofing attacks which presents a major problem in their real-life deployment. To date, most of the proposed synthetic speech detectors (SSDs) have weighted the importance of different segments…

Sound · Computer Science 2016-10-11 Ali Khodabakhsh , Cenk Demiroglu

While automatic subjective speech quality assessment has witnessed much progress, an open question is whether an automatic quality assessment at frame resolution is possible. This would be highly desirable, as it adds explainability to the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-10 Michael Kuhlmann , Fritz Seebauer , Petra Wagner , Reinhold Haeb-Umbach

Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like…

Computation and Language · Computer Science 2026-01-15 Fadhil Muhammad , Alwin Djuliansah , Adrian Aryaputra Hamzah , Kurniawati Azizah

Voice disorders significantly impact patient quality of life, yet non-invasive automated diagnosis remains under-explored due to both the scarcity of pathological voice data, and the variability in recording sources. This work introduces…

Numerous voice conversion (VC) techniques have been proposed for the conversion of voices among different speakers. Although good quality of the converted speech can be observed when VC is applied in a clean environment, the quality…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-20 Yun-Ju Chan , Chiang-Jen Peng , Syu-Siang Wang , Hsin-Min Wang , Yu Tsao , Tai-Shih Chi

Dysarthric speech recognition faces challenges from severity variations and disparities relative to normal speech. Conventional approaches individually fine-tune ASR models pre-trained on normal speech per patient to prevent feature…

Sound · Computer Science 2025-08-27 Qing Xiao , Yingshan Peng , PeiPei Zhang

Text-to-speech and co-speech gesture synthesis have until now been treated as separate areas by two different research communities, and applications merely stack the two technologies using a simple system-level pipeline. This can lead to…

Human-Computer Interaction · Computer Science 2021-08-27 Siyang Wang , Simon Alexanderson , Joakim Gustafson , Jonas Beskow , Gustav Eje Henter , Éva Székely

Advances in neural speech synthesis have brought us technology that is not only close to human naturalness, but is also capable of instant voice cloning with little data, and is highly accessible with pre-trained models available.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-03 Lauri Juvela , Xin Wang