English

Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation

Computation and Language 2026-01-06 v1 Audio and Speech Processing

Abstract

Segmenting speech transcripts into thematic sections benefits both downstream processing and users who depend on written text for accessibility. We introduce a novel approach to hierarchical topic segmentation in transcripts, generating multi-level tables of contents that capture both topic and subtopic boundaries. We compare zero-shot prompting and LoRA fine-tuning on large language models, while also exploring the integration of high-level speech pause features. Evaluations on English meeting recordings and multilingual lecture transcripts (Portuguese, German) show significant improvements over established topic segmentation baselines. Additionally, we adapt a common evaluation measure for multi-level segmentation, taking into account all hierarchical levels within one metric.

Keywords

Cite

@article{arxiv.2601.02128,
  title  = {Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation},
  author = {Steffen Freisinger and Philipp Seeberger and Thomas Ranzenberger and Tobias Bocklet and Korbinian Riedhammer},
  journal= {arXiv preprint arXiv:2601.02128},
  year   = {2026}
}

Comments

Published in Proceedings of Interspeech 2025. Please cite the proceedings version (DOI: 10.21437/Interspeech.2025-2792)

R2 v1 2026-07-01T08:50:54.252Z