English

OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL

Computation and Language 2026-06-29 v1 Machine Learning Sound Audio and Speech Processing

Abstract

We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives. OLIVE combines view-augmented masked latent prediction with waveform reconstruction under a unified objective. Reconstruction constrains early encoder features to retain signal-level information, while masked latent prediction shapes later contextual representations toward invariance for robust downstream performance. We show that these objectives enable representations that support a broad range of tasks. In particular, OLIVE improves results on generation and speaker tasks, maintains competitive performance on recognition and semantic tasks, and improves waveform reconstruction.

Cite

@article{arxiv.2606.30356,
  title  = {OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL},
  author = {Karl El Hajal and Mathew Magimai. -Doss},
  journal= {arXiv preprint arXiv:2606.30356},
  year   = {2026}
}