English

When Large Language Models Meet Speech: A Survey on Integration Approaches

Computation and Language 2025-09-10 v2 Sound Audio and Speech Processing

Abstract

Recent advancements in large language models (LLMs) have spurred interest in expanding their application beyond text-based tasks. A large number of studies have explored integrating other modalities with LLMs, notably speech modality, which is naturally related to text. This paper surveys the integration of speech with LLMs, categorizing the methodologies into three primary approaches: text-based, latent-representation-based, and audio-token-based integration. We also demonstrate how these methods are applied across various speech-related applications and highlight the challenges in this field to offer inspiration for

Keywords

Cite

@article{arxiv.2502.19548,
  title  = {When Large Language Models Meet Speech: A Survey on Integration Approaches},
  author = {Zhengdong Yang and Shuichiro Shimizu and Yahan Yu and Chenhui Chu},
  journal= {arXiv preprint arXiv:2502.19548},
  year   = {2025}
}

Comments

Accepted at Findings of ACL 2025 (Long Paper)

R2 v1 2026-06-28T21:59:19.899Z