English
Related papers

Related papers: PoDAR: Power-Disentangled Audio Representation for…

200 papers

Using unsupervised learning to disentangle speech into content, rhythm, pitch, and timbre for voice conversion has become a hot research topic. Existing works generally take into account disentangling speech components through human-crafted…

Sound · Computer Science 2024-05-01 Ziqi Liang , Jianzong Wang , Xulong Zhang , Yong Zhang , Ning Cheng , Jing Xiao

Modern speaker verification models use deep neural networks to encode utterance audio into discriminative embedding vectors. During the training process, these networks are typically optimized to differentiate arbitrary speakers. This…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-09 Hua Shen , Yuguang Yang , Guoli Sun , Ryan Langman , Eunjung Han , Jasha Droppo , Andreas Stolcke

Disentangled and interpretable latent representations in generative models typically come at the cost of generation quality. The $\beta$-VAE framework introduces a hyperparameter $\beta$ to balance disentanglement and reconstruction…

Machine Learning · Computer Science 2025-07-10 Anshuk Uppal , Yuhta Takida , Chieh-Hsin Lai , Yuki Mitsufuji

In this paper, we propose a variational autoencoder with disentanglement priors, VAE-DPRIOR, for task-specific natural language generation with none or a handful of task-specific labeled examples. In order to tackle compositional…

Computation and Language · Computer Science 2022-11-01 Zhuang Li , Lizhen Qu , Qiongkai Xu , Tongtong Wu , Tianyang Zhan , Gholamreza Haffari

Generating long-form 44.1kHz stereo audio from text prompts can be computationally demanding. Further, most previous works do not tackle that music and sound effects naturally vary in their duration. Our research focuses on the efficient…

Sound · Computer Science 2024-05-14 Zach Evans , CJ Carr , Josiah Taylor , Scott H. Hawley , Jordi Pons

We propose an algorithm, guided variational autoencoder (Guided-VAE), that is able to learn a controllable generative model by performing latent representation disentanglement learning. The learning objective is achieved by providing…

Computer Vision and Pattern Recognition · Computer Science 2020-04-06 Zheng Ding , Yifan Xu , Weijian Xu , Gaurav Parmar , Yang Yang , Max Welling , Zhuowen Tu

In order to build language technologies for majority of the languages, it is important to leverage the resources available in public domain on the internet - commonly referred to as `Found Data'. However, such data is characterized by the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-27 Nishant Gurunath , Sai Krishna Rallabandi , Alan Black

Diffusion models have achieved state-of-the-art synthesis quality on both visual and audio tasks, and recent works further adapt them to textual data by diffusing on the embedding space. In this paper, we conduct systematic studies of the…

Computation and Language · Computer Science 2024-04-23 Zhujin Gao , Junliang Guo , Xu Tan , Yongxin Zhu , Fang Zhang , Jiang Bian , Linli Xu

Disentangled sequential autoencoders (DSAEs) represent a class of probabilistic graphical models that describes an observed sequence with dynamic latent variables and a static latent variable. The former encode information at a frame rate…

Sound · Computer Science 2022-06-16 Yin-Jyun Luo , Sebastian Ewert , Simon Dixon

Generative models have thrived in computer vision, enabling unprecedented image processes. Yet the results in audio remain less advanced. Our project targets real-time sound synthesis from a reduced set of high-level parameters, including…

Sound · Computer Science 2019-06-25 Adrien Bitton , Philippe Esling , Antoine Caillon , Martin Fouilleul

This paper investigates continuous representations of steering vectors over frequency and microphone/source positions for augmented listening (e.g., spatial filtering and binaural rendering), enabling user-parameterized control of the…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-17 Diego Di Carlo , Shoichi Koyama , Nugraha Aditya Arie , Fontaine Mathieu , Bando Yoshiaki , Yoshii Kazuyoshi

One of the fundamental representation learning tasks is unsupervised sequential disentanglement, where latent codes of inputs are decomposed to a single static factor and a sequence of dynamic factors. To extract this latent information,…

Machine Learning · Computer Science 2025-10-09 Nimrod Berman , Ilan Naiman , Idan Arbiv , Gal Fadlon , Omri Azencot

Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs. A central reason is spectral bias, the tendency of neural networks to learn low-frequency structure before high-frequency (HF) details, which both…

Sound · Computer Science 2025-11-27 Ido Nitzan HIdekel , Gal lifshitz , Khen Cohen , Dan Raviv

Recent neural text-to-speech (TTS) models with fine-grained latent features enable precise control of the prosody of synthesized speech. Such models typically incorporate a fine-grained variational autoencoder (VAE) structure, extracting…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-11 Guangzhi Sun , Yu Zhang , Ron J. Weiss , Yuan Cao , Heiga Zen , Andrew Rosenberg , Bhuvana Ramabhadran , Yonghui Wu

A large part of the literature on learning disentangled representations focuses on variational autoencoders (VAE). Recent developments demonstrate that disentanglement cannot be obtained in a fully unsupervised setting without inductive…

Machine Learning · Computer Science 2021-02-11 Graziano Mita , Maurizio Filippone , Pietro Michiardi

Generative speech enhancement methods based on generative adversarial networks (GANs) and diffusion models have shown promising results in various speech enhancement tasks. However, their performance in very low signal-to-noise ratio (SNR)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-29 Shrishti Saha Shetu , Emanuël A. P. Habets , Andreas Brendel

This work adapts two recent architectures of generative models and evaluates their effectiveness for the conversion of whispered speech to normal speech. We incorporate the normal target speech into the training criterion of…

Disentangling complex data to its latent factors of variation is a fundamental task in representation learning. Existing work on sequential disentanglement mostly provides two factor representations, i.e., it separates the data to…

Machine Learning · Computer Science 2023-03-31 Nimrod Berman , Ilan Naiman , Omri Azencot

Diffusion probabilistic models (DPMs) have shown remarkable results on various image synthesis tasks such as text-to-image generation and image inpainting. However, compared to other generative methods like VAEs and GANs, DPMs lack a…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Yipeng Leng , Qiangjuan Huang , Zhiyuan Wang , Yangyang Liu , Haoyu Zhang

Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or…

Sound · Computer Science 2025-01-20 Shengkui Zhao , Zexu Pan , Kun Zhou , Yukun Ma , Chong Zhang , Bin Ma
‹ Prev 1 4 5 6 7 8 10 Next ›