English

VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling

Audio and Speech Processing 2026-03-06 v2 Computation and Language Machine Learning Sound Signal Processing

Abstract

Speech processing systems face a fundamental challenge: the human voice changes with age, yet few datasets support rigorous longitudinal evaluation. We introduce VoxKnesset, an open-access dataset of ~2,300 hours of Hebrew parliamentary speech spanning 2009-2025, comprising 393 speakers with recording spans of up to 15 years. Each segment includes aligned transcripts and verified demographic metadata from official parliamentary records. We benchmark modern speech embeddings (WavLM-Large, ECAPA-TDNN, Wav2Vec2-XLSR-1B) on age prediction and speaker verification under longitudinal conditions. Speaker verification EER rises from 2.15\% to 4.58\% over 15 years for the strongest model, and cross-sectionally trained age regressors fail to capture within-speaker aging, while longitudinally trained models recover a meaningful temporal signal. We publicly release the dataset and pipeline to support aging-robust speech systems and Hebrew speech processing.

Keywords

Cite

@article{arxiv.2603.01270,
  title  = {VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling},
  author = {Yanir Marmor and Arad Zulti and David Krongauz and Adam Gabet and Yoad Snapir and Yair Lifshitz and Eran Segal},
  journal= {arXiv preprint arXiv:2603.01270},
  year   = {2026}
}

Comments

4 pages, 5 figures, 2 tables

R2 v1 2026-07-01T10:58:14.646Z