English

GRASS: Unified Generation Model for Speech-to-Semantic Tasks

Computation and Language 2023-09-12 v2 Sound Audio and Speech Processing

Abstract

This paper explores the instruction fine-tuning technique for speech-to-semantic tasks by introducing a unified end-to-end (E2E) framework that generates target text conditioned on a task-related prompt for audio data. We pre-train the model using large and diverse data, where instruction-speech pairs are constructed via a text-to-speech (TTS) system. Extensive experiments demonstrate that our proposed model achieves state-of-the-art (SOTA) results on many benchmarks covering speech named entity recognition, speech sentiment analysis, speech question answering, and more, after fine-tuning. Furthermore, the proposed model achieves competitive performance in zero-shot and few-shot scenarios. To facilitate future work on instruction fine-tuning for speech-to-semantic tasks, we release our instruction dataset and code.

Keywords

Cite

@article{arxiv.2309.02780,
  title  = {GRASS: Unified Generation Model for Speech-to-Semantic Tasks},
  author = {Aobo Xia and Shuyu Lei and Yushu Yang and Xiang Guo and Hua Chai},
  journal= {arXiv preprint arXiv:2309.02780},
  year   = {2023}
}
R2 v1 2026-06-28T12:13:57.132Z