English
Related papers

Related papers: Quasi-Periodic Parallel WaveGAN: A Non-autoregress…

200 papers

Recent works of utilizing phonetic posteriograms (PPGs) for non-parallel voice conversion have significantly increased the usability of voice conversion since the source and target DBs are no longer required for matching contents. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-15 Sunghee Jung , Youngjoo Suh , Yeunju Choi , Hoirin Kim

Generative Adversarial Networks (GAN) are cutting-edge algorithms for generating new data samples based on the learned data distribution. However, its performance comes at a significant cost in terms of computation and memory requirements.…

Machine Learning · Computer Science 2022-01-25 Azzam Alhussain , Mingjie Lin

Spatiotemporal predictive learning (STPL) aims to forecast future frames from past observations and is essential across a wide range of applications. Compared with recurrent or hybrid architectures, pure convolutional models offer superior…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Xinyong Cai , Changbin Sun , Yong Wang , Hongyu Yang , Yuankai Wu

In this work, we propose WaveFlow, a small-footprint generative flow for raw audio, which is directly trained with maximum likelihood. It handles the long-range structure of 1-D waveform with a dilated 2-D convolutional architecture, while…

Sound · Computer Science 2020-06-26 Wei Ping , Kainan Peng , Kexin Zhao , Zhao Song

Gene regulatory relationships can be abstracted as a gene regulatory network (GRN), which plays a key role in characterizing complex cellular processes and pathways. Recently, graph neural networks (GNNs), as a class of deep learning…

Molecular Networks · Quantitative Biology 2023-11-07 Hui Zhang , Xuexin An , Qiang He , Yudong Yao , Yudong Zhang , Feng-Lei Fan , Yueyang Teng

In recent years, machine learning (ML) methods have become increasingly popular in wireless communication systems for several applications. A critical bottleneck for designing ML systems for wireless communications is the availability of…

Signal Processing · Electrical Eng. & Systems 2025-06-03 Satyavrat Wagle , Akshay Malhotra , Shahab Hamidi-Rad , Aditya Sant , David J. Love , Christopher G. Brinton

In recent years, machine learning (ML) methods have become increasingly popular in wireless communication systems for several applications. A critical bottleneck for designing ML systems for wireless communications is the availability of…

Signal Processing · Electrical Eng. & Systems 2025-03-28 Satyavrat Wagle , Akshay Malhotra , Shahab Hamidi-Rad , Aditya Sant , David J. Love , Christopher G. Brinton

State-of-the-art models for high-resolution image generation, such as BigGAN and VQVAE-2, require an incredible amount of compute resources and/or time (512 TPU-v3 cores) to train, putting them out of reach for the larger research…

Image and Video Processing · Electrical Eng. & Systems 2020-10-27 Seungwook Han , Akash Srivastava , Cole Hurwitz , Prasanna Sattigeri , David D. Cox

The Click-though Rate (CTR) prediction task is a basic task in recommendation system. Most of the previous researches of CTR models built based on Wide \& deep structure and gradually evolved into parallel structures with different modules.…

Machine Learning · Computer Science 2022-06-22 Ri Su , Alphonse Houssou Hounye , Cong Cao , Muzhou Hou

Communication is a crucial phase in the context of distributed training. Because parameter server (PS) frequently experiences network congestion, recent studies have found that training paradigms without a centralized server outperform the…

Optimization and Control · Mathematics 2020-12-17 Feijie Wu , Shiqi He , Yutong Yang , Haozhao Wang , Zhihao Qu , Song Guo , Weihua Zhuang

We present a deep convolutional GAN which leverages techniques from MP3/Vorbis audio compression to produce long, high-quality audio samples with long-range coherence. The model uses a Modified Discrete Cosine Transform (MDCT) data…

Sound · Computer Science 2021-01-14 Korneel van den Broek

We introduce a novel method for emotion conversion in speech that does not require parallel training data. Our approach loosely relies on a cycle-GAN schema to minimize the reconstruction error from converting back and forth between emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-12 Ravi Shankar , Jacob Sager , Archana Venkataraman

This paper presents a novel speech phase prediction model which predicts wrapped phase spectra directly from amplitude spectra by neural networks. The proposed model is a cascade of a residual convolutional network and a parallel estimation…

Sound · Computer Science 2023-02-17 Yang Ai , Zhen-Hua Ling

Our previous work, the unified source-filter GAN (uSFGAN) vocoder, introduced a novel architecture based on the source-filter theory into the parallel waveform generative adversarial network to achieve high voice quality and pitch…

Sound · Computer Science 2023-02-28 Reo Yoneyama , Yi-Chiao Wu , Tomoki Toda

Classical parametric speech coding techniques provide a compact representation for speech signals. This affords a very low transmission rate but with a reduced perceptual quality of the reconstructed signals. Recently, autoregressive deep…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-02 Ahmed Mustafa , Arijit Biswas , Christian Bergler , Julia Schottenhamml , Andreas Maier

This paper presents a waveform modeling and generation method using hierarchical recurrent neural networks (HRNN) for speech bandwidth extension (BWE). Different from conventional BWE methods which predict spectral parameters for…

Sound · Computer Science 2018-01-26 Zhen-Hua Ling , Yang Ai , Yu Gu , Li-Rong Dai

We present quasiparticle (QP) energies from fully self-consistent $GW$ (sc$GW$) calculations for a set of prototypical semiconductors and insulators within the framework of the projector-augmented wave methodology. To obtain converged…

Materials Science · Physics 2018-10-31 Manuel Grumet , Peitao Liu , Merzuk Kaltak , Jiří Klimeš , Georg Kresse

We present an approach to calculate the electronic structure for a range of materials using the quasiparticle self-consistent GW method with vertex corrections included in the screened Coulomb interaction W. This is achieved by solving the…

Materials Science · Physics 2023-10-10 Brian Cunningham , Myrta Gruening , Dimitar Pashov , Mark van Schilfgaarde

Neural waveform models such as the WaveNet are used in many recent text-to-speech systems, but the original WaveNet is quite slow in waveform generation because of its autoregressive (AR) structure. Although faster non-AR models were…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-30 Xin Wang , Shinji Takaki , Junichi Yamagishi

The recent deep generative models for static graphs that are now being actively developed have achieved significant success in areas such as molecule design. However, many real-world problems involve temporal graphs whose topology and…

Machine Learning · Computer Science 2021-03-09 Liming Zhang , Liang Zhao , Shan Qin , Dieter Pfoser