KnowThyself:一个用于大语言模型可解释性的智能助手
人工智能
2025-11-07 v1 信息检索
机器学习
多智能体系统
摘要
我们开发 KnowThyself,一款推进大语言模型 (LLM) 可解释性的智能助手。现有工具提供有用见解但仍然分散且需要编程。KnowThyself 将这些能力整合到聊天式界面中,用户可上传模型,提出自然语言问题,获取交互式可视化并获得引导性解释。在其核心, orchestrator LLM 首先改写用户查询,agent router 进一步将其导向 specialized 模块,最终输出被 contextualized 为连贯的解释。此设计降低了技术门槛,为 LLM 检查提供了一个可扩展平台。通过将整个过程嵌入对话式工作流程,KnowThyself 为可及的 LLM 可解释性提供了稳健基础。
引用
@article{arxiv.2511.03878,
title = {KnowThyself: An Agentic Assistant for LLM Interpretability},
author = {Suraj Prasai and Mengnan Du and Ying Zhang and Fan Yang},
journal= {arXiv preprint arXiv:2511.03878},
year = {2025}
}
备注
5 pages, 1 figure, Accepted for publication at the Demonstration Track of the 40th AAAI Conference on Artificial Intelligence (AAAI 26)