English
Related papers

Related papers: Continuous Autoregressive Modeling with Stochastic…

200 papers

Text-to-Speech (TTS) has recently seen great progress in synthesizing high-quality speech owing to the rapid development of parallel TTS systems, but producing speech with naturalistic prosodic variations, speaking styles and emotional…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Yinghao Aaron Li , Cong Han , Nima Mesgarani

Latent representations are at the heart of the majority of modern generative models. In the audio domain they are typically produced by a neural-audio-codec autoencoder. In this work we introduce SAME (Semantically-Aligned Music…

Sound · Computer Science 2026-05-19 Julian D. Parker , Zach Evans , CJ Carr , Zachary Zukowski , Josiah Taylor , Matthew Rice , Jordi Pons

Neural conversation models such as encoder-decoder models are easy to generate bland and generic responses. Some researchers propose to use the conditional variational autoencoder(CVAE) which maximizes the lower bound on the conditional…

Computation and Language · Computer Science 2019-11-25 Jun Gao , Wei Bi , Xiaojiang Liu , Junhui Li , Guodong Zhou , Shuming Shi

We present a new method to visualize data ensembles by constructing structured probabilistic representations in latent spaces, i.e., lower-dimensional representations of spatial data features. Our approach transforms the spatial features of…

Machine Learning · Computer Science 2025-09-17 Cenyang Wu , Qinhan Yu , Liang Zhou

The development of deep learning methods for magnetic resonance spectroscopy (MRS) is often hindered by limited availability of large, high-quality training datasets. While physics-based simulations are commonly used to mitigate this…

Standard simultaneous autoregressive (SAR) models typically assume normally distributed errors, an assumption often violated in real-world datasets that frequently exhibit non-normal, skewed, or heavy-tailed characteristics. New SAR models…

Methodology · Statistics 2025-12-16 Anjana Wijayawardhana , David Gunawan , Thomas Suesse

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, which suffer from significant information loss. Even with…

This paper proposes the continuous semantic topic embedding model (CSTEM) which finds latent topic variables in documents using continuous semantic distance function between the topics and the words by means of the variational…

Machine Learning · Statistics 2017-11-27 Namkyu Jung , Hyeong In Choi

Generative Adversarial Networks (GANs) can achieve state-of-the-art sample quality in generative modelling tasks but suffer from the mode collapse problem. Variational Autoencoders (VAE) on the other hand explicitly maximize a…

Machine Learning · Computer Science 2019-09-30 Apratim Bhattacharyya , Mario Fritz , Bernt Schiele

Variational autoencoders (VAEs) have been used extensively to discover low-dimensional latent factors governing neural activity and animal behavior. However, without careful model selection, the uncovered latent factors may reflect noise in…

Machine Learning · Computer Science 2023-12-13 Julia Huiming Wang , Dexter Tsin , Tatiana Engel

Recent efforts target spoken language models (SLMs) that not only listen but also speak for more natural human-LLM interaction. Joint speech-text modeling is a promising direction to achieve this. However, the effectiveness of recent speech…

Computation and Language · Computer Science 2026-02-06 Liang-Hsuan Tseng , Yi-Chang Chen , Kuan-Yi Lee , Da-Shan Shiu , Hung-yi Lee

Standard autoregressive language models generate text by repeatedly selecting a discrete next token, coupling prediction with irreversible commitment at every step. We show that token selection is not the only viable autoregressive…

Computation and Language · Computer Science 2026-04-07 Oshri Naparstek

Autoregressive language models are powerful and relatively easy to train. However, these models are usually trained without explicit conditioning labels and do not offer easy ways to control global aspects such as sentiment or topic during…

Computation and Language · Computer Science 2021-03-30 Tom Bosc , Pascal Vincent

Deep generative models for audio synthesis have recently been significantly improved. However, the task of modeling raw-waveforms remains a difficult problem, especially for audio waveforms and music signals. Recently, the realtime audio…

Sound · Computer Science 2022-11-17 Seokjin Lee , Minhan Kim , Seunghyeon Shin , Daeho Lee , Inseon Jang , Wootaek Lim

Several studies explore inferences based on stochastic volatility (SV) models, taking into account the stylized facts of return data. The common problem is that the latent parameters of many volatility models are high-dimensional and…

Statistical Finance · Quantitative Finance 2018-09-06 T. R. Santos

While generative methods have progressed rapidly in recent years, generating expressive prosody for an utterance remains a challenging task in text-to-speech synthesis. This is particularly true for systems that model prosody explicitly…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-02 Paul Mayer , Florian Lux , Alejandro Pérez-González-de-Martos , Angelina Elizarova , Lindsey Vanderlyn , Dirk Väth , Ngoc Thang Vu

In this paper, we describe a statistical parametric speech synthesis approach with unit-level acoustic representation. In conventional deep neural network based speech synthesis, the input text features are repeated for the entire duration…

Sound · Computer Science 2016-06-21 Sivanand Achanta , KNRK Raju Alluri , Suryakanth V Gangashetty

Token prediction stability remains a challenge in autoregressive generative models, where minor variations in early inference steps often lead to significant semantic drift over extended sequences. A structured modulation mechanism was…

State-of-the-art non-autoregressive text-to-speech (TTS) models based on FastSpeech 2 can efficiently synthesise high-fidelity and natural speech. For expressive speech datasets however, we observe characteristic audio distortions. We…

Sound · Computer Science 2024-09-19 Fabian Kögel , Bac Nguyen , Fabien Cardinaux

The recurrent neural networks (RNN) with richly distributed internal states and flexible non-linear transition functions, have overtaken the dynamic Bayesian networks such as the hidden Markov models (HMMs) in the task of modeling highly…

Machine Learning · Computer Science 2021-08-11 Jin Huang , Ming Xiao