English
Related papers

Related papers: WaveCycleGAN2: Time-domain Neural Post-filter for …

200 papers

Compared with air-conducted speech, bone-conducted speech has the unique advantage of shielding background noise. Enhancement of bone-conducted speech helps to improve its quality and intelligibility. In this paper, a novel CycleGAN with…

Sound · Computer Science 2021-11-03 Qing Pan , Teng Gao , Jian Zhou , Huabin Wang , Liang Tao , Hon Keung Kwan

Adversarial waveform generation has been a popular approach as the backend of singing voice conversion (SVC) to generate high-quality singing audio. However, the instability of GAN also leads to other problems, such as pitch jitters and U/V…

Sound · Computer Science 2022-01-26 Haohan Guo , Zhiping Zhou , Fanbo Meng , Kai Liu

Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would…

Computation and Language · Computer Science 2018-02-16 Kaizhi Qian , Yang Zhang , Shiyu Chang , Xuesong Yang , Dinei Florencio , Mark Hasegawa-Johnson

Separating two sources from an audio mixture is an important task with many applications. It is a challenging problem since only one signal channel is available for analysis. In this paper, we propose a novel framework for singing voice…

Sound · Computer Science 2017-11-15 Zhe-Cheng Fan , Yen-Lin Lai , Jyh-Shing Roger Jang

Non-autoregressive GAN-based neural vocoders are widely used due to their fast inference speed and high perceptual quality. However, they often suffer from audible artifacts such as tonal artifacts in their generated results. Therefore, we…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-11 Hyunjae Cho , Junhyeok Lee , Wonbin Jung

We propose DarkStream, a streaming speech synthesis model for real-time speaker anonymization. To improve content encoding under strict latency constraints, DarkStream combines a causal waveform encoder, a short lookahead buffer, and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-08 Waris Quamer , Ricardo Gutierrez-Osuna

This work provides a solution to the challenge of small amounts of training data in Non-Destructive Ultrasonic Testing for composite components. It was demonstrated that direct simulation alone is ineffective at producing training data that…

Image and Video Processing · Electrical Eng. & Systems 2023-11-06 Shaun McKnight , S. Gareth Pierce , Ehsan Mohseni , Christopher MacKinnon , Charles MacLeod , Tom OHare , Charalampos Loukas

Emotional Voice Conversion, or emotional VC, is a technique of converting speech from one emotion state into another one, keeping the basic linguistic information and speaker identity. Previous approaches for emotional VC need parallel data…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-09 Songxiang Liu , Yuewen Cao , Helen Meng

Thanks to the growing availability of spoofing databases and rapid advances in using them, systems for detecting voice spoofing attacks are becoming more and more capable, and error rates close to zero are being reached for the ASVspoof2015…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-05 Jaime Lorenzo-Trueba , Fuming Fang , Xin Wang , Isao Echizen , Junichi Yamagishi , Tomi Kinnunen

Video-to-speech is the process of reconstructing the audio speech from a video of a spoken utterance. Previous approaches to this task have relied on a two-step process where an intermediate representation is inferred from the video, and is…

Machine Learning · Computer Science 2022-08-17 Rodrigo Mira , Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Björn W. Schuller , Maja Pantic

The recently-developed WaveNet architecture is the current state of the art in realistic speech synthesis, consistently rated as more natural sounding for many different languages than any previous system. However, because WaveNet relies on…

Previous generative adversarial network (GAN)-based neural vocoders are trained to reconstruct the exact ground truth waveform from the paired mel-spectrogram and do not consider the one-to-many relationship of speech synthesis. This…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-11 Junhyeok Lee , Seungu Han , Hyunjae Cho , Wonbin Jung

Since the introduction of Generative Adversarial Networks (GANs) in speech synthesis, remarkable achievements have been attained. In a thorough exploration of vocoders, it has been discovered that audio waveforms can be generated at speeds…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-14 Yubing Cao , Yongming Li , Liejun Wang , Yinfeng Yu

We propose a parallel-data-free voice-conversion (VC) method that can learn a mapping from source to target speech without relying on parallel data. The proposed method is general purpose, high quality, and parallel-data free and works…

Machine Learning · Statistics 2017-12-21 Takuhiro Kaneko , Hirokazu Kameoka

This paper proposes an effective probability density distillation (PDD) algorithm for WaveNet-based parallel waveform generation (PWG) systems. Recently proposed teacher-student frameworks in the PWG system have successfully achieved a…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-29 Ryuichi Yamamoto , Eunwoo Song , Jae-Min Kim

We propose PeriodNet, a non-autoregressive (non-AR) waveform generation model with a new model structure for modeling periodic and aperiodic components in speech waveforms. The non-AR waveform generation models can generate speech waveforms…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-17 Yukiya Hono , Shinji Takaki , Kei Hashimoto , Keiichiro Oura , Yoshihiko Nankaku , Keiichi Tokuda

The speech enhancement task usually consists of removing additive noise or reverberation that partially mask spoken utterances, affecting their intelligibility. However, little attention is drawn to other, perhaps more aggressive signal…

Sound · Computer Science 2019-04-09 Santiago Pascual , Joan Serrà , Antonio Bonafonte

In recent years, various flow-based generative models have been proposed to generate high-fidelity waveforms in real-time. However, these models require either a well-trained teacher network or a number of flow steps making them…

Sound · Computer Science 2020-07-06 Hyeongju Kim , Hyeonseung Lee , Woo Hyun Kang , Sung Jun Cheon , Byoung Jin Choi , Nam Soo Kim

In this paper, we introduce MFCCGAN as a novel speech synthesizer based on adversarial learning that adopts MFCCs as input and generates raw speech waveforms. Benefiting the GAN model capabilities, it produces speech with higher…

Sound · Computer Science 2023-10-26 Mohammad Reza Hasanabadi Majid Behdad Davood Gharavian

Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous speakers across…

Sound · Computer Science 2023-02-17 Sang-gil Lee , Wei Ping , Boris Ginsburg , Bryan Catanzaro , Sungroh Yoon
‹ Prev 1 3 4 5 6 7 10 Next ›