English

SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation

Computation and Language 2026-04-21 v2 Artificial Intelligence

Abstract

Human infants, with only a few hundred hours of speech exposure, acquire basic units of new languages, highlighting a striking efficiency gap compared to the data-hungry self-supervised speech models. To address this gap, this paper introduces SpidR-Adapt for rapid adaptation of speech units to new languages using minimal unlabeled data. We cast such low-resource speech representation learning as a meta-learning problem and construct a multi-task adaptive pre-training (MAdaPT) protocol which formulates the adaptation process as a bi-level optimization framework. To enable scalable meta-training under this framework, we propose a novel heuristic solution, first-order bi-level optimization (FOBLO), avoiding heavy computation costs. Finally, we stabilize meta-training by using a robust initialization through interleaved supervision which alternates self-supervised and supervised objectives. Empirically, SpidR-Adapt achieves rapid gains in phonemic discriminability (ABX) and downstream spoken language modeling scores (sWUGGY, sBLIMP, tSC), surpassing in-domain toplines after training on less than 1h of target-language audio and delivering 100×100\times greater data efficiency than standard multi-task training. These findings highlight a practical, architecture-agnostic path toward biologically inspired, data-efficient representations. We open-source the training code and model checkpoints at https://github.com/facebookresearch/spidr-adapt.

Keywords

Cite

@article{arxiv.2512.21204,
  title  = {SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation},
  author = {Mahi Luthra and Jiayi Shen and Maxime Poli and Angelo Ortiz and Yosuke Higuchi and Youssef Benchekroun and Martin Gleize and Charles-Eric Saint-James and Dongyan Lin and Phillip Rust and Angel Villar and Surya Parimi and Vanessa Stark and Rashel Moritz and Juan Pino and Yann LeCun and Emmanuel Dupoux},
  journal= {arXiv preprint arXiv:2512.21204},
  year   = {2026}
}
R2 v1 2026-07-01T08:39:58.313Z