English
Related papers

Related papers: Wavehax: Aliasing-Free Neural Waveform Synthesis B…

200 papers

We propose Parallel WaveGAN, a distillation-free, fast, and small-footprint waveform generation method using a generative adversarial network. In the proposed method, a non-autoregressive WaveNet is trained by jointly optimizing…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-07 Ryuichi Yamamoto , Eunwoo Song , Jae-Min Kim

Variational Autoencoders (VAEs) are powerful generative models capable of learning compact latent representations. However, conventional VAEs often generate relatively blurry images due to their assumption of an isotropic Gaussian latent…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Andrew Kiruluta

We present the Neural Waveshaping Unit (NEWT): a novel, lightweight, fully causal approach to neural audio synthesis which operates directly in the waveform domain, with an accompanying optimisation (FastNEWT) for efficient CPU inference.…

Sound · Computer Science 2021-07-28 Ben Hayes , Charalampos Saitis , György Fazekas

Complex-valued processing has brought deep learning-based speech enhancement and signal extraction to a new level. Typically, the process is based on a time-frequency (TF) mask which is applied to a noisy spectrogram, while complex masks…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-02 Hendrik Schröter , Alberto N. Escalante-B. , Tobias Rosenkranz , Andreas Maier

Though widely used in image classification, convolutional neural networks (CNNs) are prone to noise interruptions, i.e. the CNN output can be drastically changed by small image noise. To improve the noise robustness, we try to integrate…

Computer Vision and Pattern Recognition · Computer Science 2021-09-01 Qiufu Li , Linlin Shen , Sheng Guo , Zhihui Lai

Large-scale, weakly-supervised speech recognition models, such as Whisper, have demonstrated impressive results on speech recognition across domains and languages. However, their application to long audio transcription via buffered or…

Sound · Computer Science 2023-07-12 Max Bain , Jaesung Huh , Tengda Han , Andrew Zisserman

Recent developments in generative models have shown that deep learning combined with traditional digital signal processing (DSP) techniques could successfully generate convincing violin samples [1], that source-excitation combined with…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Ollie McCarthy , Zohaib Ahmed

The goal of the Acoustic Question Answering (AQA) task is to answer a free-form text question about the content of an acoustic scene. It was inspired by the Visual Question Answering (VQA) task. In this paper, based on the previously…

Computation and Language · Computer Science 2024-01-15 Jerome Abdelnour , Jean Rouat , Giampiero Salvi

Deepfake speech detection presents a growing challenge as generative audio technologies continue to advance. We propose a hybrid training framework that advances detection performance through novel augmentation strategies. First, we…

Sound · Computer Science 2025-11-14 Inbal Rimon , Oren Gal , Haim Permuter

In this work, we propose DiffWave, a versatile diffusion probabilistic model for conditional and unconditional waveform generation. The model is non-autoregressive, and converts the white noise signal into structured waveform through a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-01 Zhifeng Kong , Wei Ping , Jiaji Huang , Kexin Zhao , Bryan Catanzaro

While significant advances have been made with respect to the separation of overlapping speech signals, studies have been largely constrained to mixtures of clean, near anechoic speech, not representative of many real-world scenarios.…

Sound · Computer Science 2020-02-17 Matthew Maciejewski , Gordon Wichern , Emmett McQuinn , Jonathan Le Roux

Recently, GAN based speech synthesis methods, such as MelGAN, have become very popular. Compared to conventional autoregressive based methods, parallel structures based generators make waveform generation process fast and stable. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-25 Qiao Tian , Yi Chen , Zewang Zhang , Heng Lu , Linghui Chen , Lei Xie , Shan Liu

Non-autoregressive GAN-based neural vocoders are widely used due to their fast inference speed and high perceptual quality. However, they often suffer from audible artifacts such as tonal artifacts in their generated results. Therefore, we…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-11 Hyunjae Cho , Junhyeok Lee , Wonbin Jung

Modern speech synthesis uses neural vocoders to model raw waveform samples directly. This increased versatility has expanded the scope of vocoders from speech to other domains, such as music. We address another interesting domain of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-22 Rhythm Bhatia , Tomi H. Kinnunen

Since Text-to-Speech systems typically don't produce waveforms directly, recent spoof detection studies use resynthesized waveforms from vocoders and neural audio codecs to simulate an attacker. Unlike vocoders, which are specifically…

Sound · Computer Science 2026-02-19 Yixuan Xiao , Florian Lux , Alejandro Pérez-González-de-Martos , Ngoc Thang Vu

Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using advanced machine…

Audio and Speech Processing · Electrical Eng. & Systems 2018-04-10 Xin Wang , Jaime Lorenzo-Trueba , Shinji Takaki , Lauri Juvela , Junichi Yamagishi

This paper presents a denoising and dereverberation hierarchical neural vocoder (DNR-HiNet) to convert noisy and reverberant acoustic features into a clean speech waveform. We implement it mainly by modifying the amplitude spectrum…

Sound · Computer Science 2020-11-10 Yang Ai , Haoyu Li , Xin Wang , Junichi Yamagishi , Zhenhua Ling

Whispered speech lacks vocal fold vibration and fundamental frequency, resulting in degraded acoustic cues and making whisper-to-normal (W2N) conversion challenging, especially with limited parallel data. We propose WhispEar, a…

Sound · Computer Science 2026-03-10 Zihao Fang , Yingda Shen , Zifan Guan , Tongtong Song , Zhenyi Liu , Zhizheng Wu

The performance of most speaker diarization systems with x-vector embeddings is both vulnerable to noisy environments and lacks domain robustness. Earlier work on speaker diarization using generative adversarial network (GAN) with an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-21 Monisankha Pal , Manoj Kumar , Raghuveer Peri , Tae Jin Park , So Hyun Kim , Catherine Lord , Somer Bishop , Shrikanth Narayanan

Achieving nuanced and accurate emulation of human voice has been a longstanding goal in artificial intelligence. Although significant progress has been made in recent years, the mainstream of speech synthesis models still relies on…

Sound · Computer Science 2024-03-04 Weiwei Lin , Chenhang He , Man-Wai Mak , Jiachen Lian , Kong Aik Lee
‹ Prev 1 4 5 6 7 8 10 Next ›