中文
相关论文

相关论文: Quasi-Periodic Parallel WaveGAN: A Non-autoregress…

200 篇论文

The generation of synthetic data with distributions that faithfully emulate the underlying data-generating mechanism holds paramount significance. Wasserstein Generative Adversarial Networks (WGANs) have emerged as a prominent tool for this…

机器学习 · 统计学 2025-01-08 Wenhui Sophia Lu , Chenyang Zhong , Wing Hung Wong

In this paper, we propose a novel behavior model for wideband PAs using a real-valued time-delay convolutional neural network (RVTDCNN). The input data of the model are sorted and arranged as the graph composed of the in-phase and…

信号处理 · 电气工程与系统科学 2020-05-21 Xin Hu , Zhijun Liu , Xiaofei Yu , Yulong Zhao , Wenhua Chen , Biao Hu , Xuekun Du , Xiang Li , Mohamed Helaoui , Weidong Wang , Fadhel M. Ghannouchi

This paper introduces a unified source-filter network with a harmonic-plus-noise source excitation generation mechanism. In our previous work, we proposed unified Source-Filter GAN (uSFGAN) for developing a high-fidelity neural vocoder with…

声音 · 计算机科学 2022-07-04 Reo Yoneyama , Yi-Chiao Wu , Tomoki Toda

We propose a temporally coherent generative model addressing the super-resolution problem for fluid flows. Our work represents a first approach to synthesize four-dimensional physics fields with neural networks. Based on a conditional…

机器学习 · 计算机科学 2025-03-20 You Xie , Aleksandra Franz , Mengyu Chu , Nils Thuerey

Clinical time series are often irregularly sampled, with varying sensor frequencies, missing observations, and misaligned timestamps. Prior approaches typically address these irregularities by interpolating data into regular sequences,…

机器学习 · 计算机科学 2025-12-18 Arash Hajisafi , Maria Despoina Siampou , Bita Azarijoo , Zhen Xiong , Cyrus Shahabi

We present an unsupervised non-parallel many-to-many voice conversion (VC) method using a generative adversarial network (GAN) called StarGAN v2. Using a combination of adversarial source classifier loss and perceptual loss, our model…

声音 · 计算机科学 2021-07-26 Yinghao Aaron Li , Ali Zare , Nima Mesgarani

Whispered speech lacks vocal fold vibration and fundamental frequency, resulting in degraded acoustic cues and making whisper-to-normal (W2N) conversion challenging, especially with limited parallel data. We propose WhispEar, a…

声音 · 计算机科学 2026-03-10 Zihao Fang , Yingda Shen , Zifan Guan , Tongtong Song , Zhenyi Liu , Zhizheng Wu

In the generator of typical Generative Adversarial Networks (GANs), a noise is inputted to generate fake samples via a series of convolutional operations. However, current noise generation models merely relies on the information from the…

机器学习 · 计算机科学 2020-05-15 Shaoning Zeng , Bob Zhang

In this work, we propose a new solution for parallel wave generation by WaveNet. In contrast to parallel WaveNet (van den Oord et al., 2018), we distill a Gaussian inverse autoregressive flow from the autoregressive WaveNet by minimizing a…

计算与语言 · 计算机科学 2019-02-25 Wei Ping , Kainan Peng , Jitong Chen

Remote photoplethysmography (rPPG) is a non-contact technique for measuring cardiac signals from facial videos. High-quality rPPG pulse signals are urgently demanded in many fields, such as health monitoring and emotion recognition.…

图像与视频处理 · 电气工程与系统科学 2020-06-05 Rencheng Song , Huan Chen , Juan Cheng , Chang Li , Yu Liu , Xun Chen

This paper proposes a method that allows non-parallel many-to-many voice conversion (VC) by using a variant of a generative adversarial network (GAN) called StarGAN. Our method, which we call StarGAN-VC, is noteworthy in that it (1)…

声音 · 计算机科学 2018-07-02 Hirokazu Kameoka , Takuhiro Kaneko , Kou Tanaka , Nobukatsu Hojo

Product quantisation (PQ) is a classical method for scalable vector encoding, yet it has seen limited usage for latent representations in high-fidelity image generation. In this work, we introduce PQGAN, a quantised image autoencoder that…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Denis Zavadski , Nikita Philip Tatsch , Carsten Rother

The generative adversarial networks (GANs) have facilitated the development of speech enhancement recently. Nevertheless, the performance advantage is still limited when compared with state-of-the-art models. In this paper, we propose a…

声音 · 计算机科学 2020-06-16 Andong Li , Chengshi Zheng , Renhua Peng , Cunhang Fan , Xiaodong Li

Time-frequency (TF) representations provide powerful and intuitive features for the analysis of time series such as audio. But still, generative modeling of audio in the TF domain is a subtle matter. Consequently, neural audio synthesis…

声音 · 计算机科学 2019-05-17 Andrés Marafioti , Nicki Holighaus , Nathanaël Perraudin , Piotr Majdak

The high harmonic generation in periodically corrugated submicrometer waveguides is studied numerically. Plasmonic field enhancement in the vicinity of the corrugations allows to use low pump intensities. Simultaneously, periodic placement…

光学 · 物理学 2017-07-12 Anton Husakou

Entertainment-oriented singing voice synthesis (SVS) requires a vocoder to generate high-fidelity (e.g. 48kHz) audio. However, most text-to-speech (TTS) vocoders cannot reconstruct the waveform well in this scenario. In this paper, we…

音频与语音处理 · 电气工程与系统科学 2023-09-19 Chunhui Wang , Chang Zeng , Jun Chen , Xing He

A new scheme for quasi-phasematching high harmonic generation (HHG) in gases is proposed. In this, the rapid variation of the axial intensity resulting from excitation of more than one mode of a waveguide is used to achieve quasi…

等离子体物理 · 物理学 2009-11-13 B. Dromey , M. Zepf , M. Landreman , S. M. Hooker

Resting-state EEG offers a non-invasive view of spontaneous brain activity, yet the extraction of meaningful patterns is often constrained by limited availability of high-quality data, and heavy reliance on manually engineered EEG features.…

神经元与认知 · 定量生物学 2025-12-01 Yeganeh Farahzadi , Morteza Ansarinia , Zoltan Kekecs

Traditional voice conversion methods rely on parallel recordings of multiple speakers pronouncing the same sentences. For real-world applications however, parallel data is rarely available. We propose MelGAN-VC, a voice conversion method…

音频与语音处理 · 电气工程与系统科学 2019-12-06 Marco Pasini

Diffusion models have exhibited exciting capabilities in generating images and are also very promising for video creation. However, the inference speed of diffusion models is limited by the slow sampling process, restricting its use cases.…

计算机视觉与模式识别 · 计算机科学 2024-12-05 XiuYu Zhang , Zening Luo , Michelle E. Lu