中文

能力坐标:面向LLM评估的统一MTMM几何框架

计算与语言 2026-05-15 v2

摘要

大语言模型(LLM)的评估面临一个关键挑战,即构件有效性问题,其中碎片化的基准和临时性指标常常混淆方法方差(如提示敏感性)与真实隐含能力。同时,新兴的研究表明,LLM的能力和输出可以建模为连续几何流形。在本系统知识(SoK)中,我们通过提出面向LLM评估的通用多特征多方法(MTMM)框架来桥接这些范式。我们对九个评估指标进行形式化和统一,包括改写不稳定性、漂移得分、Overton宽度和多样性得分,将其解释为不再是孤立标量值,而是作为共享隐含坐标空间中的几何测量。这种空间统一性将模型行为分解为三个正交的隐含维度:(1)不稳定性和敏感性,(2)位置和对齐,(3)覆盖范围和表达性。通过系统性地分离与真实能力范围无关的扰动,框架为稳健且经验上稳定的基准设计提供了理论依据和通用性分类框架。

关键词

引用

@article{arxiv.2605.08522,
  title  = {Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation},
  author = {Adib Sakhawat and Tahsin Islam and Takia Farhin and Syed Rifat Raiyan and Hasan Mahmud and Md Kamrul Hasan},
  journal= {arXiv preprint arXiv:2605.08522},
  year   = {2026}
}

备注

The paper has mistake of undertaking political spaces to semantic dimensions. This needs to be removed because this is a fetal flaw in consideration. The initial hypothesis and premise needs to be rigorously formulated within the political landscape not generalizing the metrics. Hence a withdrawal for now is necessary