English

PromptSep: Generative Audio Separation via Multimodal Prompting

Sound 2025-11-07 v1 Audio and Speech Processing

Abstract

Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based approaches. However, two key limitations restrict their practical use: (1) users often require operations beyond separation, such as sound removal; and (2) relying solely on text prompts can be unintuitive for specifying sound sources. In this paper, we propose PromptSep to extend LASS into a broader framework for general-purpose sound separation. PromptSep leverages a conditional diffusion model enhanced with elaborated data simulation to enable both audio extraction and sound removal. To move beyond text-only queries, we incorporate vocal imitation as an additional and more intuitive conditioning modality for our model, by incorporating Sketch2Sound as a data augmentation strategy. Both objective and subjective evaluations on multiple benchmarks demonstrate that PromptSep achieves state-of-the-art performance in sound removal and vocal-imitation-guided source separation, while maintaining competitive results on language-queried source separation.

Keywords

Cite

@article{arxiv.2511.04623,
  title  = {PromptSep: Generative Audio Separation via Multimodal Prompting},
  author = {Yutong Wen and Ke Chen and Prem Seetharaman and Oriol Nieto and Jiaqi Su and Rithesh Kumar and Minje Kim and Paris Smaragdis and Zeyu Jin and Justin Salamon},
  journal= {arXiv preprint arXiv:2511.04623},
  year   = {2025}
}

Comments

Submitted to ICASSP 2026

R2 v1 2026-07-01T07:25:00.247Z