中文

FOReCAst:面向未来结果推理与置信度评估的基准

机器学习 2025-05-19 v4 计算与语言

摘要

预测是许多领域的重要任务,如技术与经济学。然而,现有预测基准缺乏全面的置信度评估、仅限于有限的 question 类型,且常由不符合实际人类预测需求的人工 question 组成。为填补这些差距,本文引入 FOReCAst (Future Outcome Reasoning and Confidence Assessment),以评估模型对预测及其置信度的能力。FOReCAst 涵盖多样化的预测情景,包括布尔型 question、时间范围预测与数量估计,实现对实际应用中预测准确性与置信度校准的全面评估。

关键词

引用

@article{arxiv.2502.19676,
  title  = {FOReCAst: The Future Outcome Reasoning and Confidence Assessment Benchmark},
  author = {Zhangdie Yuan and Zifeng Ding and Andreas Vlachos},
  journal= {arXiv preprint arXiv:2502.19676},
  year   = {2025}
}