English

Spontaneous Informal Speech Dataset for Punctuation Restoration

Computation and Language 2024-09-18 v1 Human-Computer Interaction Machine Learning Sound Audio and Speech Processing

Abstract

Presently, punctuation restoration models are evaluated almost solely on well-structured, scripted corpora. On the other hand, real-world ASR systems and post-processing pipelines typically apply towards spontaneous speech with significant irregularities, stutters, and deviations from perfect grammar. To address this discrepancy, we introduce SponSpeech, a punctuation restoration dataset derived from informal speech sources, which includes punctuation and casing information. In addition to publicly releasing the dataset, we contribute a filtering pipeline that can be used to generate more data. Our filtering pipeline examines the quality of both speech audio and transcription text. We also carefully construct a ``challenging" test set, aimed at evaluating models' ability to leverage audio information to predict otherwise grammatically ambiguous punctuation. SponSpeech is available at https://github.com/GitHubAccountAnonymous/PR, along with all code for dataset building and model runs.

Keywords

Cite

@article{arxiv.2409.11241,
  title  = {Spontaneous Informal Speech Dataset for Punctuation Restoration},
  author = {Xing Yi Liu and Homayoon Beigi},
  journal= {arXiv preprint arXiv:2409.11241},
  year   = {2024}
}

Comments

8 pages, 7 tables, 1 figure, Recognition Technologies, Inc. Technical Report

R2 v1 2026-06-28T18:47:54.423Z