NLPStatTest:用于比较NLP系统性能的工具包
计算与语言
2020-12-17 v1 应用统计
摘要
以p值为中心的统计显著性检验常用于比较NLP系统性能,但仅p值并不充分,因为统计显著性与实际显著性不同。后者可通过估计效应量来度量。本文提出一种比较NLP系统性能的三阶段流程,并提供自动化该过程的工具包NLPStatTest。用户可上传NLP系统评估分数,工具包将分析这些分数、运行适当的显著性检验、估计效应量,并进行功效分析以估计II类错误。该工具包提供了一种便捷且系统化的方式比较NLP系统性能,超越了统计显著性检验。
引用
@article{arxiv.2011.13231,
title = {NLPStatTest: A Toolkit for Comparing NLP System Performance},
author = {Haotian Zhu and Denise Mak and Jesse Gioannini and Fei Xia},
journal= {arXiv preprint arXiv:2011.13231},
year = {2020}
}
备注
Will appear in AACL-IJCNLP 2020