English

AMSS-Net: Audio Manipulation on User-Specified Sources with Textual Queries

Audio and Speech Processing 2021-04-29 v1 Machine Learning Sound

Abstract

This paper proposes a neural network that performs audio transformations to user-specified sources (e.g., vocals) of a given audio track according to a given description while preserving other sources not mentioned in the description. Audio Manipulation on a Specific Source (AMSS) is challenging because a sound object (i.e., a waveform sample or frequency bin) is `transparent'; it usually carries information from multiple sources, in contrast to a pixel in an image. To address this challenging problem, we propose AMSS-Net, which extracts latent sources and selectively manipulates them while preserving irrelevant sources. We also propose an evaluation benchmark for several AMSS tasks, and we show that AMSS-Net outperforms baselines on several AMSS tasks via objective metrics and empirical verification.

Keywords

Cite

@article{arxiv.2104.13553,
  title  = {AMSS-Net: Audio Manipulation on User-Specified Sources with Textual Queries},
  author = {Woosung Choi and Minseok Kim and Marco A. Martínez Ramírez and Jaehwa Chung and Soonyoung Jung},
  journal= {arXiv preprint arXiv:2104.13553},
  year   = {2021}
}

Comments

10 pages, 8 figures, 3 tables, under reviewing of ACMMM 21

R2 v1 2026-06-24T01:35:12.747Z