LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment
Abstract
Fueled by increasing model scale and multimodal inputs, Multimodal Large Language Models (MLLMs) have emerged as a promising paradigm for Spoken Language Assessment (SLA). While effective, this paradigm often overlooks the intrinsic ordinal structure of language acquisition. This paper works around the necessity of large-scale MLLMs by introducing Latent Ordinal Prototype Alignment (LOPA) for SLA, a prototype-based regularizer that enforces an ordinal geometric prior directly on the latent space. Coupled with Semantic-Anchored Layer Routing (SALR), which adaptively harvests multi-depth representations from a frozen Whisper encoder, our framework achieves an RMSE of 0.361. This performance rivals billion-parameter systems without the need for LLM-based fine-tuning. Further analysis reveals that SALR's synergy with LOPA offers interpretable, criterion-aligned preferences, thereby supporting an efficient and ordinal-aware modeling alternative to current scaling-centric models for SLA.
Keywords
Cite
@article{arxiv.2606.31310,
title = {LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment},
author = {Hong-Yun Lin and Fu-An Chao and Bi-Cheng Yan and Berlin Chen},
journal= {arXiv preprint arXiv:2606.31310},
year = {2026}
}