English
Related papers

Related papers: MaskCycleGAN-based Whisper to Normal Speech Conver…

200 papers

In this paper, we present a description of the baseline system of Voice Conversion Challenge (VCC) 2020 with a cyclic variational autoencoder (CycleVAE) and Parallel WaveGAN (PWG), i.e., CycleVAEPWG. CycleVAE is a nonparallel VAE-based…

Sound · Computer Science 2020-10-12 Patrick Lumban Tobing , Yi-Chiao Wu , Tomoki Toda

Whisper, as a form of speech, is not sufficiently addressed by mainstream speech applications. This is due to the fact that systems built for normal speech do not work as expected for whispered speech. A first step to building a speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-27 S. Johanan Joysingh , P. Vijayalakshmi , T. Nagarajan

Numerous models have shown great success in the fields of speech recognition as well as speech synthesis, but models for speech to speech processing have not been heavily explored. We propose Speech to Speech Synthesis Network (STSSN), a…

Sound · Computer Science 2026-02-20 Bjorn Johnson , Jared Levy

Deep learning models have improved sign language-to-text translation and made it easier for non-signers to understand signed messages. When the goal is spoken communication, a naive approach is to convert signed messages into text and then…

Sound · Computer Science 2026-04-14 Toranosuke Manabe , Yuto Shibata , Shinnosuke Takamichi , Yoshimitsu Aoki

This paper proposes a framework for modeling sound change that combines deep learning and iterative learning. Acquisition and transmission of speech is modeled by training generations of Generative Adversarial Networks (GANs) on unannotated…

Computation and Language · Computer Science 2021-09-23 Gašper Beguš

Generative adversarial networks have seen rapid development in recent years and have led to remarkable improvements in generative modelling of images. However, their application in the audio domain has received limited attention, and…

The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our study applies a modified Transformer in a speech enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-04 Szu-Wei Fu , Chien-Feng Liao , Tsun-An Hsieh , Kuo-Hsuan Hung , Syu-Siang Wang , Cheng Yu , Heng-Cheng Kuo , Ryandhimas E. Zezario , You-Jin Li , Shang-Yi Chuang , Yen-Ju Lu , Yu Tsao

Voice profiling aims at inferring various human parameters from their speech, e.g. gender, age, etc. In this paper, we address the challenge posed by a subtask of voice profiling - reconstructing someone's face from their voice. The task is…

Sound · Computer Science 2019-06-04 Yandong Wen , Rita Singh , Bhiksha Raj

Benefiting from massive and diverse data sources, speech foundation models exhibit strong generalization and knowledge transfer capabilities to a wide range of downstream tasks. However, a limitation arises from their exclusive handling of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-10 Pengcheng Guo , Xuankai Chang , Hang Lv , Shinji Watanabe , Lei Xie

Machine Speech Chain, simulating the human perception-production loop, proves effective in jointly improving ASR and TTS. We propose TokenChain, a fully discrete speech chain coupling semantic-token ASR with a two-stage TTS: an…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-05 Mingxuan Wang , Satoshi Nakamura

Automatic speech recognition systems have created exciting possibilities for applications, however they also enable opportunities for systematic eavesdropping. We propose a method to camouflage a person's voice over-the-air from these…

Sound · Computer Science 2022-02-18 Mia Chiquier , Chengzhi Mao , Carl Vondrick

Recently, neural vocoders have been widely used in speech synthesis tasks, including text-to-speech and voice conversion. However, when encountering data distribution mismatch between training and inference, neural vocoders trained on real…

Sound · Computer Science 2020-08-21 Po-chun Hsu , Chun-hsuan Wang , Andy T. Liu , Hung-yi Lee

Large self-supervised pre-trained speech models have achieved remarkable success across various speech-processing tasks. The self-supervised training of these models leads to universal speech representations that can be used for different…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-25 Vamsikrishna Chemudupati , Marzieh Tahaei , Heitor Guimaraes , Arthur Pimentel , Anderson Avila , Mehdi Rezagholizadeh , Boxing Chen , Tiago Falk

We propose a novel method that combines CycleGAN and inter-domain losses for semi-supervised end-to-end automatic speech recognition. Inter-domain loss targets the extraction of an intermediate shared representation of speech and text…

Computation and Language · Computer Science 2022-10-24 Chia-Yu Li , Ngoc Thang Vu

This research introduces an enhanced version of the multi-objective speech assessment model--MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the effectiveness of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-30 Ryandhimas E. Zezario , Yu-Wen Chen , Szu-Wei Fu , Yu Tsao , Hsin-Min Wang , Chiou-Shann Fuh

Voice conversion is to generate a new speech with the source content and a target voice style. In this paper, we focus on one general setting, i.e., non-parallel many-to-many voice conversion, which is close to the real-world scenario. As…

Sound · Computer Science 2022-07-28 Jian Ma , Zhedong Zheng , Hao Fei , Feng Zheng , Tat-seng Chua , Yi Yang

Adversarial loss in a conditional generative adversarial network (GAN) is not designed to directly optimize evaluation metrics of a target task, and thus, may not always guide the generator in a GAN to generate data with improved metric…

Sound · Computer Science 2019-05-14 Szu-Wei Fu , Chien-Feng Liao , Yu Tsao , Shou-De Lin

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-17 Alexander Polok , Dominik Klement , Matthew Wiesner , Sanjeev Khudanpur , Jan Černocký , Lukáš Burget

Despite the remarkable progress made in synthesizing emotional speech from text, it is still challenging to provide emotion information to existing speech segments. Previous methods mainly rely on parallel data, and few works have studied…

Sound · Computer Science 2020-03-06 Xiaoqi Jia , Jianwei Tai , Hang Zhou , Yakai Li , Weijuan Zhang , Haichao Du , Qingjia Huang

Previous works (Donahue et al., 2018a; Engel et al., 2019a) have found that generating coherent raw audio waveforms with GANs is challenging. In this paper, we show that it is possible to train GANs reliably to generate high quality…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-10 Kundan Kumar , Rithesh Kumar , Thibault de Boissiere , Lucas Gestin , Wei Zhen Teoh , Jose Sotelo , Alexandre de Brebisson , Yoshua Bengio , Aaron Courville
‹ Prev 1 4 5 6 7 8 10 Next ›