AI 基准测试结果分析:关键指标及其获取方法
人工智能
2019-03-25 v2
摘要
项目反应理论(IRT)可用于分析 AI 基准测试结果。两参数 IRT 模型在项目(或 AI 问题)一侧提供两个指标(难度与区分度),而在被试(或 AI 智能体)一侧仅提供一个指标(能力)。本文我们分析如何通过在被试一侧增加第四个指标——通用性(generality),使这组指标成为对偶的。通用性意在与区分度对偶,并基于难度。即,通用性被定义为一种新度量,用于评估一个智能体是否稳定地擅长简单问题而不擅长困难问题。加入通用性后,我们看到这组四个关键指标能让我们对 AI 基准测试结果获得更深入的认识。特别地,我们考察了 AI 中两个流行的基准:Arcade Learning Environment(Atari 2600 游戏)与 General Video Game AI 竞赛。我们给出了针对其他 AI 基准与竞赛估计并解释这些指标的若干准则。
引用
@article{arxiv.1811.08186,
title = {Analysing Results from AI Benchmarks: Key Indicators and How to Obtain Them},
author = {Fernando Martínez-Plumed and José Hernández-Orallo},
journal= {arXiv preprint arXiv:1811.08186},
year = {2019}
}
备注
This report is a preliminary version of a related paper with title "Dual Indicators to Analyse AI Benchmarks: Difficulty, Discrimination, Ability and Generality", accepted for publication at IEEE Transactions on Games. Please refer to and cite the journal paper (https://doi.org/10.1109/TG.2018.2883773)