English

Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition

Computation and Language 2026-01-13 v1

Abstract

In speech language modeling, two architectures dominate the frontier: the Transformer and the Conformer. However, it remains unknown whether their comparable performance stems from convergent processing strategies or distinct architectural inductive biases. We introduce Architectural Fingerprinting, a probing framework that isolates the effect of architecture on representation, and apply it to a controlled suite of 24 pre-trained encoders (39M-3.3B parameters). Our analysis reveals divergent hierarchies: Conformers implement a "Categorize Early" strategy, resolving phoneme categories 29% earlier in depth and speaker gender by 16% depth. In contrast, Transformers "Integrate Late," deferring phoneme, accent, and duration encoding to deep layers (49-57%). These fingerprints suggest design heuristics: Conformers' front-loaded categorization may benefit low-latency streaming, while Transformers' deep integration may favor tasks requiring rich context and cross-utterance normalization.

Keywords

Cite

@article{arxiv.2601.06972,
  title  = {Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition},
  author = {Nathan Roll and Pranav Bhalerao and Martijn Bartelds and Arjun Pawar and Yuka Tatsumi and Tolulope Ogunremi and Chen Shani and Calbert Graham and Meghan Sumner and Dan Jurafsky},
  journal= {arXiv preprint arXiv:2601.06972},
  year   = {2026}
}

Comments

3 figures, 9 tables

R2 v1 2026-07-01T08:59:40.736Z