Omni-R1:微调音频 LLM 是否真的需要音频?
音频与语音处理
2025-11-24 v3 声音
摘要
我们提出 Omni-R1,通过强化学习方法 GRPO 对最新多模态 LLM Qwen2.5-Omni 进行微调,使其在音频问答数据集上训练。该方法在最新的 MMAU 和 MMAR 基准测试中实现了新的最佳成绩。Omni-R1 在 sounds、music、speech 和 overall average 类别中,分别在 Test-mini 和 Test-full 数据集上取得最高准确率。为理解性能提升,我们测试了带和不带音频的模型,发现 GRPO 带来的大部分性能提升可归因于更好的基于文本的推理。我们还发现了一个意外结果:在文本数据集上进行无音频微调,仍能有效提升音频-based 性能。
引用
@article{arxiv.2505.09439,
title = {Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?},
author = {Andrew Rouditchenko and Saurabhchand Bhati and Edson Araujo and Samuel Thomas and Hilde Kuehne and Rogerio Feris and James Glass},
journal= {arXiv preprint arXiv:2505.09439},
year = {2025}
}