中文
相关论文

相关论文: Wavetable Synthesis Using CVAE for Timbre Control …

200 篇论文

In this study, a deep learning based conditional density estimation technique known as conditional variational auto-encoder (CVAE) is used to fill gaps typically observed in particle image velocimetry (PIV) measurements in combustion…

流体动力学 · 物理学 2023-12-12 Shashank Yellapantula

Timbre is a primary mode of expression in diverse musical contexts. However, prevalent audio-driven synthesis methods predominantly rely on pitch and loudness envelopes, effectively flattening timbral expression from the input. Our approach…

声音 · 计算机科学 2024-07-08 Jordie Shier , Charalampos Saitis , Andrew Robertson , Andrew McPherson

This paper proposes a new model, called condition-transforming variational autoencoder (CTVAE), to improve the performance of conversation response generation using conditional variational autoencoders (CVAEs). In conventional CVAEs , the…

计算与语言 · 计算机科学 2019-04-25 Yu-Ping Ruan , Zhen-Hua Ling , Quan Liu , Zhigang Chen , Nitin Indurkhya

Recently, several deep learning methods are proposed for the gravitational wave data analysis. One is conditional variational auto encoder (CVAE), proposed by Gabbard et al. [1]. We study the accuracy of a CVAE in the context of the…

广义相对论与量子宇宙学 · 物理学 2020-02-28 Takahiro S. Yamamoto , Takahiro Tanaka

Bioacoustics data from Passive acoustic monitoring (PAM) poses a unique set of challenges for classification, particularly the limited availability of complete and reliable labels in datasets due to annotation uncertainty, biological…

The generation of virtual populations (VPs) of anatomy is essential for conducting in silico trials of medical devices. Typically, the generated VP should capture sufficient variability while remaining plausible and should reflect the…

图像与视频处理 · 电气工程与系统科学 2023-07-31 Haoran Dou , Nishant Ravikumar , Alejandro F. Frangi

Voice conversion is a task of synthesizing an utterance with target speaker's voice while maintaining linguistic information of the source utterance. While a speaker can produce varying utterances from a single script with different…

声音 · 计算机科学 2025-04-17 Soobin Suh , Dabi Ahn , Heewoong Park , Jonghun Park

Neural networks are used for channel decoding, channel detection, channel evaluation, and resource management in multi-input and multi-output (MIMO) wireless communication systems. In this paper, we consider the problem of finding precoding…

信号处理 · 电气工程与系统科学 2022-05-06 Evgeny Bobrov , Alexander Markov , Sviatoslav Panchenko , Dmitry Vetrov

Although the Conditional Variational AutoEncoder (CVAE) model can generate more diversified responses than the traditional Seq2Seq model, the responses often have low relevance with the input words or are illogical with the question. A…

计算与语言 · 计算机科学 2022-10-12 Jiayi Liu , Wei Wei , Zhixuan Chu , Xing Gao , Ji Zhang , Tan Yan , Yulin Kang

Deep generative models for audio synthesis have recently been significantly improved. However, the task of modeling raw-waveforms remains a difficult problem, especially for audio waveforms and music signals. Recently, the realtime audio…

声音 · 计算机科学 2022-11-17 Seokjin Lee , Minhan Kim , Seunghyeon Shin , Daeho Lee , Inseon Jang , Wootaek Lim

Variational autoencoders (VAEs) are widely used deep generative models capable of learning unsupervised latent representations of data. Such representations are often difficult to interpret or control. We consider the problem of…

机器学习 · 计算机科学 2018-12-18 Jack Klys , Jake Snell , Richard Zemel

Re-orchestration is the process of adapting a music piece for a different set of instruments. By altering the original instrumentation, the orchestrator often modifies the musical texture while preserving a recognizable melodic line and…

声音 · 计算机科学 2025-07-01 Dinh-Viet-Toan Le , Yi-Hsuan Yang

The surprisingness of a song is an essential and seemingly subjective factor in determining whether the listener likes it. With the help of information theory, it can be described as the transition probability of a music sequence modeled as…

声音 · 计算机科学 2021-08-25 Yi-Wei Chen , Hung-Shin Lee , Yen-Hsing Chen , Hsin-Min Wang

We present a conditional variational autoencoder (CVAE) that generates stellar spectra covering 4000 $\le$ $T_{\mathrm{eff}$ $\le$ 11,000 K, $2.0 \le \log g \le 5.0$ dex, $-1.5 \le [\mathrm{M}/\mathrm{H}] \le +1.5$ dex, $v\sin i \le 300$…

太阳与恒星天体物理 · 物理学 2025-08-26 Marwan Gebran , Ian Bentley

Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesised speech of a target speaker's timbre. Most previous approaches rely on data with style labels, but manually-annotated labels are…

声音 · 计算机科学 2022-12-14 Chunyu Qiang , Peng Yang , Hao Che , Xiaorui Wang , Zhongyuan Wang

We investigate large-scale latent variable models (LVMs) for neural story generation -- an under-explored application for open-domain long text -- with objectives in two threads: generation effectiveness and controllability. LVMs,…

计算与语言 · 计算机科学 2021-07-09 Le Fang , Tao Zeng , Chaochun Liu , Liefeng Bo , Wen Dong , Changyou Chen

Controllable data generation aims to synthesize data by specifying values for target concepts. Achieving this reliably requires modeling the underlying generative factors and their relationships. In real-world scenarios, these factors…

机器学习 · 计算机科学 2025-11-21 Qilong Zhao , Shiyu Wang , Zeeshan Memon , Yang Qiao , Guangji Bai , Bo Pan , Zhaohui Qin , Liang Zhao

Syntactic information contains structures and rules about how text sentences are arranged. Incorporating syntax into text modeling methods can potentially benefit both representation learning and generation. Variational autoencoders (VAEs)…

计算与语言 · 计算机科学 2019-08-28 Yijun Xiao , William Yang Wang

In this paper we present a new implementation of a Variational Autoencoder (VAE) for the calibration of sensors. We propose that the VAE can be used to calibrate sensor data by training the latent space as a calibration output. We discuss…

机器学习 · 计算机科学 2025-11-04 Travis Barrett , Amit Kumar Mishra , Joyce Mwangama

Neural audio synthesis methods can achieve high-fidelity and realistic sound generation by utilizing deep generative models. Such models typically rely on external labels which are often discrete as conditioning information to achieve…

声音 · 计算机科学 2024-06-12 Yunyi Liu , Craig Jin