English

SALM: Speech-augmented Language Model with In-context Learning for Speech Recognition and Translation

Computation and Language 2023-10-17 v1 Human-Computer Interaction Sound Audio and Speech Processing

Abstract

We present a novel Speech Augmented Language Model (SALM) with {\em multitask} and {\em in-context} learning capabilities. SALM comprises a frozen text LLM, a audio encoder, a modality adapter module, and LoRA layers to accommodate speech input and associated task instructions. The unified SALM not only achieves performance on par with task-specific Conformer baselines for Automatic Speech Recognition (ASR) and Speech Translation (AST), but also exhibits zero-shot in-context learning capabilities, demonstrated through keyword-boosting task for ASR and AST. Moreover, {\em speech supervised in-context training} is proposed to bridge the gap between LLM training and downstream speech tasks, which further boosts the in-context learning ability of speech-to-text models. Proposed model is open-sourced via NeMo toolkit.

Keywords

Cite

@article{arxiv.2310.09424,
  title  = {SALM: Speech-augmented Language Model with In-context Learning for Speech Recognition and Translation},
  author = {Zhehuai Chen and He Huang and Andrei Andrusenko and Oleksii Hrinchuk and Krishna C. Puvvada and Jason Li and Subhankar Ghosh and Jagadeesh Balam and Boris Ginsburg},
  journal= {arXiv preprint arXiv:2310.09424},
  year   = {2023}
}

Comments

submit to ICASSP 2024

R2 v1 2026-06-28T12:50:24.889Z