中文
相关论文

相关论文: Universal audio synthesizer control with normalizi…

200 篇论文

Voice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, these approaches lack…

声音 · 计算机科学 2021-02-02 Paarth Neekhara , Shehzeen Hussain , Shlomo Dubnov , Farinaz Koushanfar , Julian McAuley

Despite significant advancements in deep learning for vision and natural language, unsupervised domain adaptation in audio remains relatively unexplored. We, in part, attribute this to the lack of an appropriate benchmark dataset. To…

声音 · 计算机科学 2023-09-27 Chia-Hsin Lin , Charles Jones , Björn W. Schuller , Harry Coppock

While self-supervised learning (SSL) has revolutionized audio representation, the excessive parameterization and quadratic computational cost of standard Transformers limit their deployment on resource-constrained devices. To address this…

声音 · 计算机科学 2026-03-30 Harunori Kawano , Takeshi Sasaki

This paper proposes a novel way of doing audio synthesis at the waveform level using Transformer architectures. We propose a deep neural network for generating waveforms, similar to wavenet. This is fully probabilistic, auto-regressive, and…

声音 · 计算机科学 2021-07-09 Prateek Verma , Chris Chafe

This paper develops a novel control synthesis approach for a wide class of practical systems. The control action is derived by inserting a compensator device in the forward path of the system that is to be controlled. The compensator design…

系统与控制 · 电气工程与系统科学 2023-02-28 Ahmad A. Masoud

Granular sound synthesis is a popular audio generation technique based on rearranging sequences of small waveform windows. In order to control the synthesis, all grains in a given corpus are analyzed through a set of acoustic descriptors.…

声音 · 计算机科学 2021-07-06 Adrien Bitton , Philippe Esling , Tatsuya Harada

We propose an algorithm that is capable of synthesizing high quality target speaker's singing voice given only their normal speech samples. The proposed algorithm first integrate speech and singing synthesis into a unified framework, and…

声音 · 计算机科学 2019-12-24 Liqiang Zhang , Chengzhu Yu , Heng Lu , Chao Weng , Yusong Wu , Xiang Xie , Zijin Li , Dong Yu

Can uncorrelated surrounding sound sources be used to generate extended diffuse sound fields? By definition, targets are a constant sound pressure level, a vanishing average sound intensity, uncorrelated sound waves arriving isotropically…

音频与语音处理 · 电气工程与系统科学 2024-02-22 Franz Zotter , Stefan Riedel , Lukas Gölles , Matthias Frank

This paper describes a human-in-the-loop approach to personalized voice synthesis in the absence of reference speech data from the target speaker. It is intended to help vocally disabled individuals restore their lost voices without…

音频与语音处理 · 电气工程与系统科学 2025-05-27 Yusheng Tian , Junbin Liu , Tan Lee

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events with their mixture and…

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse…

Recent progress in diffusion-based audio generation and restoration has substantially improved performance across heterogeneous conditioning regimes, including text-conditioned audio generation and audio-conditioned super-resolution.…

声音 · 计算机科学 2026-05-07 Xuanhao Zhang , Chang Li

Controlling audible sound requires inherently broadband and subwavelength acoustic solutions, which are to date, crucially missing. This includes current noise absorption methods, such as porous materials or acoustic resonators, which are…

应用物理 · 物理学 2023-05-26 Stanislav Sergeev , Hervé Lissek , Romain Fleury

We consider the problem of learning high-level controls over the global structure of generated sequences, particularly in the context of symbolic music generation with complex language models. In this work, we present the Transformer…

声音 · 计算机科学 2020-07-01 Kristy Choi , Curtis Hawthorne , Ian Simon , Monica Dinculescu , Jesse Engel

Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker identity within a single utterance. This underexplored…

Data-driven controller design based on data informativity has gained popularity due to its straightforward applicability, while providing rigorous guarantees. However, applying this framework to the estimator synthesis problem introduces…

系统与控制 · 电气工程与系统科学 2025-04-14 Felix Brändle , Frank Allgöwer

Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to increase the generalizability, explainability, and fairness…

音频与语音处理 · 电气工程与系统科学 2023-08-09 Michael Kuhlmann , Adrian Meise , Fritz Seebauer , Petra Wagner , Reinhold Haeb-Umbach

Audio autoencoders learn useful, compressed audio representations, but their non-linear latent spaces prevent intuitive algebraic manipulation such as mixing or scaling. We introduce a simple training methodology to induce linearity in a…

声音 · 计算机科学 2026-01-29 Bernardo Torres , Manuel Moussallam , Gabriel Meseguer-Brocal

Analogy-making is a key method for computer algorithms to generate both natural and creative music pieces. In general, an analogy is made by partially transferring the music abstractions, i.e., high-level representations and their…

声音 · 计算机科学 2019-10-22 Ruihan Yang , Dingsu Wang , Ziyu Wang , Tianyao Chen , Junyan Jiang , Gus Xia

Speaker generation task aims to create unseen speaker voice without reference speech. The key to the task is defining a speaker space that represents diverse speakers to determine the generated speaker trait. However, the effective way to…

声音 · 计算机科学 2025-07-08 Masato Murata , Koichi Miyazaki , Tomoki Koriyama , Tomoki Toda
‹ 上一页 1 8 9 10 下一页 ›