中文
相关论文

相关论文: Real-time Timbre Remapping with Differentiable DSP

200 篇论文

Controllable text-to-speech (TTS) systems face significant challenges in achieving independent manipulation of speaker timbre and speaking style, often suffering from entanglement between these attributes. We present DMP-TTS, a latent…

声音 · 计算机科学 2025-12-11 Kang Yin , Chunyu Qiang , Sirui Zhao , Xiaopeng Wang , Yuzhe Liang , Pengfei Cai , Tong Xu , Chen Zhang , Enhong Chen

Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment…

声音 · 计算机科学 2026-05-06 Jan Melechovsky , Ambuj Mehrish , Abhinaba Roy , Dorien Herremans

Multiple moving sound source localization in real-world scenarios remains a challenging issue due to interaction between sources, time-varying trajectories, distorted spatial cues, etc. In this work, we propose to use deep learning…

声音 · 计算机科学 2022-02-17 Bing Yang , Hong Liu , Xiaofei Li

Understanding how large audio models represent music, and using that understanding to steer generation, is both challenging and underexplored. Inspired by mechanistic interpretability in language models, where direction vectors in…

Text-to-music generation has advanced rapidly, with modern autoregressive and diffusion-based models producing convincing music from natural-language prompts. However, much of this progress relies on large-scale training data and external…

声音 · 计算机科学 2026-05-21 Junyoung Koh

The ubiquity of sound synthesizers has reshaped music production and even entirely defined new music genres. However, the increasing complexity and number of parameters in modern synthesizers make them harder to master. Hence, the…

机器学习 · 计算机科学 2019-07-03 Philippe Esling , Naotake Masuda , Adrien Bardet , Romeo Despres , Axel Chemla--Romeu-Santos

Modern digital music production typically involves combining numerous acoustic elements to compile a piece of music. Important types of such elements are drum samples, which determine the characteristics of the percussive components of the…

声音 · 计算机科学 2022-08-03 Stefan Lattner

Although digital cameras can acquire high-dynamic range (HDR) images, the captured HDR information are mostly quantized to low-dynamic range (LDR) images for display compatibility and compact storage. In this paper, we propose an invertible…

图像与视频处理 · 电气工程与系统科学 2021-10-12 Zhuming Zhang , Menghan Xia , Xueting Liu , Chengze Li , Tien-Tsin Wong

We propose a transformer-based rhythm quantization model that incorporates beat and downbeat information to quantize MIDI performances into metrically-aligned, human-readable scores. We propose a beat-based preprocessing method that…

声音 · 计算机科学 2025-08-28 Maximilian Wachter , Sebastian Murgul , Michael Heizmann

Choral music separation refers to the task of extracting tracks of voice parts (e.g., soprano, alto, tenor, and bass) from mixed audio. The lack of datasets has impeded research on this topic as previous work has only been able to train and…

Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules,…

声音 · 计算机科学 2025-10-13 Zhao Guo , Ziqian Ning , Guobin Ma , Lei Xie

Guitar tablature is a form of music notation widely used among guitarists. It captures not only the musical content of a piece, but also its implementation and ornamentation on the instrument. Guitar Tablature Transcription (GTT) is an…

声音 · 计算机科学 2025-05-27 Yongyi Zang , Yi Zhong , Frank Cwitkowitz , Zhiyao Duan

The advancement of diffusion-based text-to-music generation has opened new avenues for zero-shot music editing. However, existing methods fail to achieve stem-specific timbre transfer, which requires altering specific stems while strictly…

声音 · 计算机科学 2026-05-12 Haowen Li , Tianxiang Li , Yi Yang , Boyu Cao , Qi Liu

Motivated by the state-of-art psychological research, we note that a piano performance transcribed with existing Automatic Music Transcription (AMT) methods cannot be successfully resynthesized without affecting the artistic content of the…

声音 · 计算机科学 2026-01-21 Federico Simonetta , Stavros Ntalampiras , Federico Avanzini

Systems for synthesizer sound matching, which automatically set the parameters of a synthesizer to emulate an input sound, have the potential to make the process of synthesizer programming faster and easier for novice and experienced…

音频与语音处理 · 电气工程与系统科学 2024-07-24 Fred Bruford , Frederik Blang , Shahan Nercessian

In this paper, a sparse-based method for the estimation of the parameters of multidimensional ($R$-D) modal (harmonic or damped) complex signals in noise is presented. The problem is formulated as $R$ simultaneous sparse approximations of…

信息论 · 计算机科学 2015-11-02 Souleymen Sahnoun , El-Hadi Djermoune , David Brie , Pierre Comon

We propose a novel Neural Steering technique that adapts the target area of a spatial-aware multi-microphone sound source separation algorithm during inference without the necessity of retraining the deep neural network (DNN). To achieve…

音频与语音处理 · 电气工程与系统科学 2024-10-23 Martin Strauss , Wolfgang Mack , María Luis Valero , Okan Köpüklü

Music source separation is the task of separating a mixture of instruments into constituent tracks. Music source separation models are typically trained using only audio data, although additional information can be used to improve the…

音频与语音处理 · 电气工程与系统科学 2025-06-04 Eetu Tunturi , David Diaz-Guerra , Archontis Politis , Tuomas Virtanen

Score Distillation Sampling (SDS) is a recent but already widely popular method that relies on an image diffusion model to control optimization problems using text prompts. In this paper, we conduct an in-depth analysis of the SDS loss…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Thiemo Alldieck , Nikos Kolotouros , Cristian Sminchisescu

The recently developed pitch-controllable text-to-speech (TTS) model, i.e. FastPitch, was conditioned for the pitch contours. However, the quality of the synthesized speech degraded considerably for pitch values that deviated significantly…

音频与语音处理 · 电气工程与系统科学 2022-04-13 Hanbin Bae , Young-Sun Joo