中文

KnowThyself:一个用于大语言模型可解释性的智能助手

人工智能 2025-11-07 v1 信息检索 机器学习 多智能体系统

摘要

我们开发 KnowThyself,一款推进大语言模型 (LLM) 可解释性的智能助手。现有工具提供有用见解但仍然分散且需要编程。KnowThyself 将这些能力整合到聊天式界面中,用户可上传模型,提出自然语言问题,获取交互式可视化并获得引导性解释。在其核心, orchestrator LLM 首先改写用户查询,agent router 进一步将其导向 specialized 模块,最终输出被 contextualized 为连贯的解释。此设计降低了技术门槛,为 LLM 检查提供了一个可扩展平台。通过将整个过程嵌入对话式工作流程,KnowThyself 为可及的 LLM 可解释性提供了稳健基础。

关键词

引用

@article{arxiv.2511.03878,
  title  = {KnowThyself: An Agentic Assistant for LLM Interpretability},
  author = {Suraj Prasai and Mengnan Du and Ying Zhang and Fan Yang},
  journal= {arXiv preprint arXiv:2511.03878},
  year   = {2025}
}

备注

5 pages, 1 figure, Accepted for publication at the Demonstration Track of the 40th AAAI Conference on Artificial Intelligence (AAAI 26)