中文

TISDiSS:面向判别式源分离的训练时与推理时可扩展框架

声音 2025-10-15 v3 人工智能 音频与语音处理

摘要

源分离是语音、音乐和音频处理中的基础任务,也为训练生成模型提供更干净、更多的数据。然而,实际中提升分离性能往往依赖于日益增大的网络,导致训练和部署成本的不断上升。近期在生成模型推理时扩展方面取得的进展启发我们,提出Training-Time and Inference-Time Scalable Discriminative Source Separation (TISDiSS),一个统一的框架,集成早期分裂多损失监督、共享参数设计和动态推理重复机制。TISDiSS通过调整推理深度即可实现灵活的速度与性能权衡,无需重新训练额外模型。我们进一步对架构和训练选择进行系统性分析,表明更多的推理重复有助于改善浅层推理性能,受益于低延迟应用。在标准语音分离基准测试上进行实验,表明该方法在降低参数数量的同时实现了最先进的性能,确立了TISDiSS作为可扩展且实用的自适应源分离框架。代码已公开于 https://github.com/WingSingFung/TISDiSS。

关键词

引用

@article{arxiv.2509.15666,
  title  = {TISDiSS: A Training-Time and Inference-Time Scalable Framework for Discriminative Source Separation},
  author = {Yongsheng Feng and Yuetonghui Xu and Jiehui Luo and Hongjia Liu and Xiaobing Li and Feng Yu and Wei Li},
  journal= {arXiv preprint arXiv:2509.15666},
  year   = {2025}
}

备注

Submitted to ICASSP 2026.(C) 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work