English
Related papers

Related papers: Deep Feature CycleGANs: Speaker Identity Preservin…

200 papers

Recently, cycle-consistent adversarial network (Cycle-GAN) has been successfully applied to voice conversion to a different speaker without parallel data, although in those approaches an individual model is needed for each target speaker.…

Audio and Speech Processing · Electrical Eng. & Systems 2018-06-26 Ju-chieh Chou , Cheng-chieh Yeh , Hung-yi Lee , Lin-shan Lee

Real-world audio recordings are often degraded by factors such as noise, reverberation, and equalization distortion. This paper introduces HiFi-GAN, a deep learning method to transform recorded speech to sound as though it had been recorded…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-23 Jiaqi Su , Zeyu Jin , Adam Finkelstein

Nowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available. Although a speaker-dependent (SD) vocoder usually outperforms a speaker-independent (SI) vocoder, it is impractical to collect a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-11 Yi-Chiao Wu , Cheng-Hung Hu , Hung-Shin Lee , Yu-Huai Peng , Wen-Chin Huang , Yu Tsao , Hsin-Min Wang , Tomoki Toda

Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-17 Shota Horiguchi , Takanori Ashihara , Marc Delcroix , Atsushi Ando , Naohiro Tawara

Deep learning models have improved sign language-to-text translation and made it easier for non-signers to understand signed messages. When the goal is spoken communication, a naive approach is to convert signed messages into text and then…

Sound · Computer Science 2026-04-14 Toranosuke Manabe , Yuto Shibata , Shinnosuke Takamichi , Yoshimitsu Aoki

Speaker adaptation methods aim to create fair quality synthesis speech voice font for target speakers while only limited resources available. Recently, as deep neural networks based statistical parametric speech synthesis (SPSS) methods…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-08 Zhiying Huang , Heng Lu , Ming Lei , Zhijie Yan

The scarcity of speaker-annotated far-field speech presents a significant challenge in developing high-performance far-field speaker verification (SV) systems. While data augmentation using large-scale near-field speech has been a common…

Sound · Computer Science 2025-01-16 Li Zhang , Jiyao Liu , Lei Xie

In this paper, we address the problem of speaker verification in conditions unseen or unknown during development. A standard method for speaker verification consists of extracting speaker embeddings with a deep neural network and processing…

Sound · Computer Science 2021-08-18 Luciana Ferrer , Mitchell McLaren , Niko Brummer

Non-parallel voice conversion (VC) is a technique for training voice converters without a parallel corpus. Cycle-consistent adversarial network-based VCs (CycleGAN-VC and CycleGAN-VC2) are widely accepted as benchmark methods. However,…

Sound · Computer Science 2021-02-26 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Nobukatsu Hojo

This paper focuses on using voice conversion (VC) to improve the speech intelligibility of surgical patients who have had parts of their articulators removed. Due to the difficulty of data collection, VC without parallel data is highly…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-26 Li-Wei Chen , Hung-Yi Lee , Yu Tsao

This paper aims to improve the widely used deep speaker embedding x-vector model. We propose the following improvements: (1) a hybrid neural network structure using both time delay neural network (TDNN) and long short-term memory neural…

Computation and Language · Computer Science 2019-02-22 Yun Tang , Guohong Ding , Jing Huang , Xiaodong He , Bowen Zhou

Modeling the rich prosodic variations inherent in human speech is essential for generating natural-sounding speech. While speaker embeddings are commonly used as conditioning inputs in personalized speech generation, they are typically…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-22 Ismail Rasim Ulgen , John H. L. Hansen , Carlos Busso , Berrak Sisman

This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker embeddings to enhance the performance of the phrasing model. We…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Dong Yang , Yuki Saito , Takaaki Saeki , Tomoki Koriyama , Wataru Nakata , Detai Xin , Hiroshi Saruwatari

Performance degradation caused by language mismatch is a common problem when applying a speaker verification system on speech data in different languages. This paper proposes a domain transfer network, named EDITnet, to alleviate the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-16 Jingyu Li , Wei Liu , Tan Lee

This paper is focused on the finetuning of acoustic models for speaker adaptation goals on a given gender. We pretrained the Transformer baseline model on Librispeech-960 and conduct experiments with finetuning on the gender-specific test…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-18 Sokolov Artem , Andrey V. Savchenko

Most of the current deep learning-based approaches for speech enhancement only operate in the spectrogram or waveform domain. Although a cross-domain transformer combining waveform- and spectrogram-domain inputs has been proposed, its…

Sound · Computer Science 2023-10-31 Jialu Li , Junhui Li , Pu Wang , Youshan Zhang

While deep learning models have demonstrated robust performance in speaker recognition tasks, they primarily rely on low-level audio features learned empirically from spectrograms or raw waveforms. However, prior work has indicated that…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-10 Nick Mehlman , Thomas Thebaud , Dani Byrd , Shri Narayanan

Robust speaker recognition, including in the presence of malicious attacks, is becoming increasingly important and essential, especially due to the proliferation of several smart speakers and personal agents that interact with an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-19 Arindam Jati , Chin-Cheng Hsu , Monisankha Pal , Raghuveer Peri , Wael AbdAlmageed , Shrikanth Narayanan

This paper proposes a deep neural network (DNN)-based multi-channel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal. DNN-based multi-channel speech enhancement is often…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-17 Yoshiki Masuyama , Masahito Togami , Tatsuya Komatsu

As is expressed in the adage "a picture is worth a thousand words", when using spoken language to communicate visual information, brevity can be a challenge. This work describes a novel technique for leveraging machine-learned feature…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-27 Andrew Port , Chelhwon Kim , Mitesh Patel