InsightBench:通过多步骤洞察生成评估业务分析智能体
摘要
数据分析对于从数据中提取有价值洞察以帮助组织作出有效决策至关重要。我们介绍 InsightBench,一个具有三个关键特征的基准数据集。首先,它由 100 个代表金融、事件管理等多样业务场景的数据集组成,每个数据集都伴随一组精心策划的植入式洞察。其次,与专注于回答单个查询的现有基准不同,InsightBench 根据智能体在端到端数据分析中的能力进行评估,包括制定问题、解释答案以及生成洞察和可操作步骤的总结。此外,我们对基准数据集进行了全面的质量保证,确保每个数据集都有清晰目标,并包含相关且有意义的问题和分析。 Furthermore, we implement a two-way evaluation mechanism using LLaMA-3 as an effective, open-source evaluator to assess agents' ability to extract insights. We also propose AgentPoirot, our baseline data analysis agent capable of performing end-to-end data analytics. Our evaluation on InsightBench shows that AgentPoirot outperforms existing approaches (such as Pandas Agent) that focus on resolving single queries. We also compare the performance of open- and closed-source LLMs and various evaluation strategies. Overall, this benchmark serves as a testbed to motivate further development in comprehensive automated data analytics and can be accessed here: https://github.com/ServiceNow/insight-bench.
引用
@article{arxiv.2407.06423,
title = {InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation},
author = {Gaurav Sahu and Abhay Puri and Juan Rodriguez and Amirhossein Abaskohi and Mohammad Chegini and Alexandre Drouin and Perouz Taslakian and Valentina Zantedeschi and Alexandre Lacoste and David Vazquez and Nicolas Chapados and Christopher Pal and Sai Rajeswar Mudumba and Issam Hadj Laradji},
journal= {arXiv preprint arXiv:2407.06423},
year = {2025}
}
备注
Accepted to ICLR 2025