English
Related papers

Related papers: Accent Conversion: A Problem-Driven Survey of Soci…

200 papers

We propose autoencoding speaker conversion for training data augmentation in automatic speech translation. This technique directly transforms an audio sequence, resulting in audio synthesized to resemble another speaker's voice. Our method…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-28 Arya D. McCarthy , Liezl Puzon , Juan Pino

In this paper, we demonstrate the efficacy of transfer learning and continuous learning for various automatic speech recognition (ASR) tasks. We start with a pre-trained English ASR model and show that transfer learning can be effectively…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-12 Jocelyn Huang , Oleksii Kuchaiev , Patrick O'Neill , Vitaly Lavrukhin , Jason Li , Adriana Flores , Georg Kucsko , Boris Ginsburg

The last decade has witnessed significant advancements in deep learning-based speech enhancement (SE). However, most existing SE research has limitations on the coverage of SE sub-tasks, data diversity and amount, and evaluation metrics. To…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-25 Wangyou Zhang , Robin Scheibler , Kohei Saijo , Samuele Cornell , Chenda Li , Zhaoheng Ni , Anurag Kumar , Jan Pirklbauer , Marvin Sach , Shinji Watanabe , Tim Fingscheidt , Yanmin Qian

While signal conversion and disentangled representation learning have shown promise for manipulating data attributes across domains such as audio, image, and multimodal generation, existing approaches, especially for speech style…

Sound · Computer Science 2025-10-10 Jonathan Svirsky , Ofir Lindenbaum , Uri Shaham

Voice disorders are pathologies significantly affecting patient quality of life. However, non-invasive automated diagnosis of these pathologies is still under-explored, due to both a shortage of pathological voice data, and diversity of the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Alkis Koudounas , Gabriele Ciravegna , Marco Fantini , Giovanni Succo , Erika Crosetti , Tania Cerquitelli , Elena Baralis

Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. Similar alignment has been reported in Whisper-style speech encoders, based on spoken translation retrieval using representational…

Computation and Language · Computer Science 2026-04-07 Ryan Soh-Eun Shim , Domenico De Cristofaro , Chengzhi Martin Hu , Alessandro Vietti , Barbara Plank

Regional accents of the same language affect not only how words are pronounced (i.e., phonetic content), but also impact prosodic aspects of speech such as speaking rate and intonation. This paper investigates a novel flow-based approach to…

Code-switching poses a number of challenges and opportunities for multilingual automatic speech recognition. In this paper, we focus on the question of robust and fair evaluation metrics. To that end, we develop a reference benchmark data…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-30 Injy Hamed , Amir Hussein , Oumnia Chellah , Shammur Chowdhury , Hamdy Mubarak , Sunayana Sitaram , Nizar Habash , Ahmed Ali

Traditional models of accent perception underestimate the role of gradient variations in phonological features which listeners rely upon for their accent judgments. We investigate how pretrained representations from current self-supervised…

Sound · Computer Science 2025-06-25 Nitin Venkateswaran , Kevin Tang , Ratree Wayland

The successful application of neural methods to machine translation has realized huge quality advances for the community. With these improvements, many have noted outstanding challenges, including the modeling and treatment of gendered…

Computation and Language · Computer Science 2020-10-16 Hila Gonen , Kellie Webster

Language in the Arab world presents a complex diglossic and multilingual setting, involving the use of Modern Standard Arabic, various dialects and sub-dialects, as well as multiple European languages. This diverse linguistic landscape has…

Computation and Language · Computer Science 2025-01-24 Injy Hamed , Caroline Sabty , Slim Abdennadher , Ngoc Thang Vu , Thamar Solorio , Nizar Habash

Human talkers often address listeners with language-comprehension challenges, such as hard-of-hearing or non-native adults, by globally slowing down their speech. However, it remains unclear whether this strategy actually makes speech more…

Computation and Language · Computer Science 2026-04-01 Paige Tuttösí , Angelica Lim , H. Henny Yeung , Yue Wang , Jean-Julien Aucouturier

When only a limited amount of accented speech data is available, to promote multi-accent speech recognition performance, the conventional approach is accent-specific adaptation, which adapts the baseline model to multiple target accents…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-06 Han Zhu , Li Wang , Pengyuan Zhang , Yonghong Yan

For many low-resource or endangered languages, spoken language resources are more likely to be annotated with translations than with transcriptions. Recent work exploits such annotations to produce speech-to-translation alignments, without…

Computation and Language · Computer Science 2017-02-16 Antonios Anastasopoulos , David Chiang

Proactive dialogue systems, related to a wide range of real-world conversational applications, equip the conversational agent with the capability of leading the conversation direction towards achieving pre-defined targets or fulfilling…

Computation and Language · Computer Science 2023-05-10 Yang Deng , Wenqiang Lei , Wai Lam , Tat-Seng Chua

Language is a cornerstone of cultural identity, yet globalization and the dominance of major languages have placed nearly 3,000 languages at risk of extinction. Existing AI-driven translation models prioritize efficiency but often fail to…

Computation and Language · Computer Science 2025-06-10 Mahfuz Ahmed Anik , Abdur Rahman , Azmine Toushik Wasi , Md Manjurul Ahsan

Researches have shown accent classification can be improved by integrating semantic information into pure acoustic approach. In this work, we combine phonetic knowledge, such as vowels, with enhanced acoustic features to build an improved…

Sound · Computer Science 2016-02-25 Zhenhao Ge

Estimating the semantic similarity between text data is one of the challenging and open research problems in the field of Natural Language Processing (NLP). The versatility of natural language makes it difficult to define rule-based methods…

Computation and Language · Computer Science 2021-02-24 Dhivya Chandrasekaran , Vijay Mago

Speech emotion conversion is the task of modifying the perceived emotion of a speech utterance while preserving the lexical content and speaker identity. In this study, we cast the problem of emotion conversion as a spoken language…

A large number of machine translation approaches have recently been developed to facilitate the fluid migration of content across languages. However, the literature suggests that many obstacles must still be dealt with to achieve better…

Computation and Language · Computer Science 2019-07-26 Diego Moussallem , Matthias Wauer , Axel-Cyrille Ngonga Ngomo