中文
相关论文

相关论文: ViolinDiff: Enhancing Expressive Violin Synthesis …

200 篇论文

Recently, diffusion models have shown remarkable results in image synthesis by gradually removing noise and amplifying signals. Although the simple generative process surprisingly works well, is this the best way to generate image data? For…

计算机视觉与模式识别 · 计算机科学 2022-11-22 Sangyun Lee , Hyungjin Chung , Jaehyeon Kim , Jong Chul Ye

Tracking the fundamental frequency (f0) of a monophonic instrumental performance is effectively a solved problem with several solutions achieving 99% accuracy. However, the related task of automatic music transcription requires a further…

声音 · 计算机科学 2023-11-16 Xavier Riley , Simon Dixon

Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address…

音频与语音处理 · 电气工程与系统科学 2025-01-17 Siyuan Hou , Shansong Liu , Ruibin Yuan , Wei Xue , Ying Shan , Mangsuo Zhao , Chao Zhang

We recently developed a neural network that receives as input the geometrical and mechanical parameters that define a violin top plate and gives as output its first ten eigenfrequencies computed in free boundary conditions. In this…

声音 · 计算机科学 2021-02-19 Davide Salvi , Sebastian Gonzalez , Fabio Antonacci , Augusto Sarti

Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms their efficiency. Alternatively, we propose VoiceFlow, an…

音频与语音处理 · 电气工程与系统科学 2024-09-04 Yiwei Guo , Chenpeng Du , Ziyang Ma , Xie Chen , Kai Yu

In music-driven dance motion generation, most existing methods use hand-crafted features and neglect that music foundation models have profoundly impacted cross-modal content generation. To bridge this gap, we propose a diffusion-based…

声音 · 计算机科学 2025-02-28 Xinran Liu , Zhenhua Feng , Diptesh Kanojia , Wenwu Wang

Recent MIDI-to-audio synthesis methods using deep neural networks have successfully generated high-quality, expressive instrumental tracks. However, these methods require MIDI annotations for supervised training, limiting the diversity of…

声音 · 计算机科学 2025-06-12 Osamu Take , Taketo Akama

The task of bandwidth extension addresses the generation of missing high frequencies of audio signals based on knowledge of the low-frequency part of the sound. This task applies to various problems, such as audio coding or audio…

声音 · 计算机科学 2023-11-28 Pierre-Amaury Grumiaux , Mathieu Lagrange

Text-to-image diffusion models have demonstrated an impressive ability to produce high-quality outputs. However, they often struggle to accurately follow fine-grained spatial information in an input text. To this end, we propose a…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Ran Galun , Sagie Benaim

Modeling and synthesizing low-light raw noise is a fundamental problem for computational photography and image processing applications. Although most recent works have adopted physics-based models to synthesize noise, the signal-independent…

计算机视觉与模式识别 · 计算机科学 2023-10-25 Feng Zhang , Bin Xu , Zhiqiang Li , Xinran Liu , Qingbo Lu , Changxin Gao , Nong Sang

Audio diffusion models can synthesize a wide variety of sounds. Existing models often operate on the latent domain with cascaded phase recovery modules to reconstruct waveform. This poses challenges when generating high-fidelity audio. In…

声音 · 计算机科学 2023-11-21 Ge Zhu , Yutong Wen , Marc-André Carbonneau , Zhiyao Duan

We introduce the Latent Fourier Transform (LatentFT), a framework that provides novel frequency-domain controls for generative music models. LatentFT combines a diffusion autoencoder with a latent-space Fourier transform to separate musical…

声音 · 计算机科学 2026-04-21 Mason Wang , Cheng-Zhi Anna Huang

We present a feature engineering pipeline for the construction of musical signal characteristics, to be used for the design of a supervised model for musical genre identification. The key idea is to extend the traditional two-step process…

声音 · 计算机科学 2021-04-08 Tina Raissi , Alessandro Tibo , Paolo Bientinesi

While most music generation models use textual or parametric conditioning (e.g. tempo, harmony, musical genre), we propose to condition a language model based music generation system with audio input. Our exploration involves two distinct…

声音 · 计算机科学 2024-07-31 Simon Rouard , Yossi Adi , Jade Copet , Axel Roebel , Alexandre Défossez

This paper presents DiffMoog - a differentiable modular synthesizer with a comprehensive set of modules typically found in commercial instruments. Being differentiable, it allows integration into neural networks, enabling automated sound…

音频与语音处理 · 电气工程与系统科学 2024-01-24 Noy Uzrad , Oren Barkan , Almog Elharar , Shlomi Shvartzman , Moshe Laufer , Lior Wolf , Noam Koenigstein

Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling vibrato remains…

声音 · 计算机科学 2025-10-07 Joon-Seung Choi , Dong-Min Byun , Hyung-Seok Oh , Seong-Whan Lee

We introduce a film score generation framework to harmonize visual pixels and music melodies utilizing a latent diffusion model. Our framework processes film clips as input and generates music that aligns with a general theme while offering…

多媒体 · 计算机科学 2024-11-13 F. Qi , L. Ni , C. Xu

Finger vein authentication, recognized for its high security and specificity, has become a focal point in biometric research. Traditional methods predominantly concentrate on vein feature extraction for discriminative modeling, with a…

计算机视觉与模式识别 · 计算机科学 2024-02-06 Yanjun Liu , Wenming Yang , Qingmin Liao

The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes…

音频与语音处理 · 电气工程与系统科学 2024-01-19 Tan Dat Nguyen , Ji-Hoon Kim , Youngjoon Jang , Jaehun Kim , Joon Son Chung

Whispered speech is characterised by a noise-like excitation that results in the lack of fundamental frequency. Considering that prosodic phenomena such as intonation are perceived through f0 variation, the perception of whispered prosody…

音频与语音处理 · 电气工程与系统科学 2023-07-07 Pablo Pérez Zarazaga , Zofia Malisz