DeepTest 工具竞赛 2026:面向基于 LLM 的汽车助手的基准测试
人工智能
2026-04-15 v1
摘要
本报告总结了首届大型语言模型 (LLM) 测试竞赛的结果,该竞赛作为 DeepTest 工作坊的一部分在 ICSE 2026 上举办。四个工具参与了对 LLM 基于汽车用户手册信息检索应用的基准测试,目标是识别系统未能恰当提及手册中警示信息的用户输入。测试解决方案依据其暴露故障的有效性以及发现的测试样本的多样性进行评估。我们报告了实验方法、参赛者情况以及测试结果。
引用
@article{arxiv.2604.12615,
title = {DeepTest Tool Competition 2026: Benchmarking an LLM-Based Automotive Assistant},
author = {Lev Sorokin and Ivan Vasilev and Samuele Pasini},
journal= {arXiv preprint arXiv:2604.12615},
year = {2026}
}
备注
Published in the proceedings of the DeepTest workshop at the 48th International Conference on Software Engineering (ICSE) 2026