English
Related papers

Related papers: A new cosine series antialiasing function and its …

200 papers

Accent normalization (AN) systems often struggle with unnatural outputs and undesired content distortion, stemming from both suboptimal training data and rigid duration modeling. In this paper, we propose a "source-synthesis" methodology…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-24 Qibing Bai , Shuhao Shi , Shuai Wang , Yukai Ju , Yannan Wang , Haizhou Li

Rapid advancements in generative modeling have made synthetic audio generation easy, making speech-based services vulnerable to spoofing attacks. Consequently, there is a dire need for robust countermeasures more than ever. Existing…

Sound · Computer Science 2025-09-03 Arnab Das , Yassine El Kheir , Carlos Franzreb , Tim Herzig , Tim Polzehl , Sebastian Möller

In this paper, we present a cross-lingual voice cloning approach. BN features obtained by SI-ASR model are used as a bridge across speakers and language boundaries. The relationships between text and BN features are modeled by the latent…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-31 Xinyong Zhou , Hao Che , Xiaorui Wang , Lei Xie

Modeling and estimation of the vocal tract and glottal source parameters of vowels from raw speech can be typically done by using the Auto-Regressive with eXogenous input (ARX) model and Liljencrants-Fant (LF) model with an iteration-based…

Sound · Computer Science 2024-10-08 Kai Lia , Masato Akagia , Yongwei Lib , Masashi Unokia

In this paper, we provide a comprehensive theory of anti-aliasing sampling patterns that explains and revises known results, and show how patterns as predicted by the theory can be generated via a variational optimization framework. We…

Graphics · Computer Science 2019-02-25 A. Cengiz Öztireli

We propose AudioStyleGAN (ASGAN), a new generative adversarial network (GAN) for unconditional speech synthesis. As in the StyleGAN family of image synthesis models, ASGAN maps sampled noise to a disentangled latent vector which is then…

Sound · Computer Science 2022-10-12 Matthew Baas , Herman Kamper

Neural networks have become ubiquitous with guitar distortion effects modelling in recent years. Despite their ability to yield perceptually convincing models, they are susceptible to frequency aliasing when driven by high frequency and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-19 Alistair Carson , Alec Wright , Stefan Bilbao

Deep generative models have achieved significant progress in speech synthesis to date, while high-fidelity singing voice synthesis is still an open problem for its long continuous pronunciation, rich high-frequency parts, and strong…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-08 Rongjie Huang , Chenye Cui , Feiyang Chen , Yi Ren , Jinglin Liu , Zhou Zhao , Baoxing Huai , Zhefeng Wang

In neural audio synthesis, neural vocoders and codecs are models that reconstruct waveforms from acoustic and latent representations, which are essential to the resulting audio quality. While current models are capable of generating…

Sound · Computer Science 2026-05-14 Yicheng Gu , Junan Zhang , Chaoren Wang , Jerry Li , Zhizheng Wu , Lauri Juvela

This paper proposes a source-filter-based generative adversarial neural vocoder named SF-GAN, which achieves high-fidelity waveform generation from input acoustic features by introducing F0-based source excitation signals to a neural filter…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-24 Ye-Xin Lu , Yang Ai , Zhen-Hua Ling

The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused…

Currently, the development of Foreign Accent Conversion (FAC) models utilizes deep neural network architectures, as well as ensembles of neural networks for speech recognition and speech generation. The use of these models is limited by…

Sound · Computer Science 2024-05-24 Vladimir Nechaev , Sergey Kosyakov

We present SURE-Score: an approach for learning score-based generative models using training samples corrupted by additive Gaussian noise. When a large training set of clean samples is available, solving inverse problems via score-based…

Machine Learning · Computer Science 2025-04-23 Asad Aali , Marius Arvinte , Sidharth Kumar , Jonathan I. Tamir

Here, we propose a new reconstruction method of smooth time-series signals. A key concept of this study is not considering the model in signal space, but in delay-embedded space. In other words, we indirectly represent a time-series signal…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-21 Tatsuya Yokota

In a recent paper, we have presented a generative adversarial network (GAN)-based model for unconditional generation of the mel-spectrograms of singing voices. As the generator of the model is designed to take a variable-length sequence of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-13 Jen-Yu Liu , Yu-Hua Chen , Yin-Cheng Yeh , Yi-Hsuan Yang

This paper introduces a unified source-filter network with a harmonic-plus-noise source excitation generation mechanism. In our previous work, we proposed unified Source-Filter GAN (uSFGAN) for developing a high-fidelity neural vocoder with…

Sound · Computer Science 2022-07-04 Reo Yoneyama , Yi-Chiao Wu , Tomoki Toda

Some glottal analysis approaches based upon linear prediction or complex cepstrum approaches have been proved to be effective to estimate glottal source from real speech utterances. We propose a new approach employing both an all-pole…

Sound · Computer Science 2016-12-16 Yiqiao Chen , John N. Gowdy

It is challenging to build a multi-singer high-fidelity singing voice synthesis system with cross-lingual ability by only using monolingual singers in the training stage. In this paper, we propose CrossSinger, which is a cross-lingual…

Sound · Computer Science 2023-09-25 Xintong Wang , Chang Zeng , Jun Chen , Chunhui Wang

Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated speech quality over…

Audio and Speech Processing · Electrical Eng. & Systems 2018-04-26 Lauri Juvela , Vassilis Tsiaras , Bajibabu Bollepalli , Manu Airaksinen , Junichi Yamagishi , Paavo Alku

Common deep learning approaches for antibody engineering focus on modeling the marginal distribution of sequences. By treating sequences as independent samples, however, these methods overlook affinity maturation as a rich and largely…

Machine Learning · Computer Science 2026-05-28 Stephen Zhewen Lu , Aakarsh Vermani , Kohei Sanno , Jiarui Lu , Frederick A Matsen , Milind Jagota , Yun S. Song
‹ Prev 1 2 3 10 Next ›