中文

大型语言模型大小对数据到文本生成中事实不一致的指数比例:统计验证

计算与语言 2025-02-19 v1 人工智能 机器学习

摘要

监控数据到文本生成(D2T)中的事实不一致性对于确保其可信度至关重要。虽然大型语言模型(LLMs)在各种 D2T 任务中显示出卓越的性能,但以前的研究主要关注通过幂律比例到 LLM 大小(即模型参数数量)的泛化误差。然而,尚无研究考察 LLM 大小对 D2T 中事实不一致性的影响。本文通过探索两种比例定律:幂律和指数比例,来调查 D2T 中事实不一致性随 LLM 大小的比例。为严格评估和比较这些比例定律,我们采用由三个关键阶段组成的统计验证框架:预测性能估计、拟合优度评估和比较分析。For a comprehensive empirical study, we analyze three popular LLM families across five D2T datasets, measuring factual inconsistency inversely using four state-of-the-art consistency metrics. Our findings, based on exhaustive empirical results and validated through our framework, reveal that, contrary to the widely assumed power law scaling, factual inconsistency in D2T follows an exponential scaling with LLM size.

关键词

引用

@article{arxiv.2502.12372,
  title  = {Factual Inconsistency in Data-to-Text Generation Scales Exponentially with LLM Size: A Statistical Validation},
  author = {Joy Mahapatra and Soumyajit Roy and Utpal Garain},
  journal= {arXiv preprint arXiv:2502.12372},
  year   = {2025}
}

备注

21 pages