中文

从 LLM 推理到自主 AI 代理:全面综述

人工智能 2026-03-10 v2 机器学习

摘要

大型语言模型和自主 AI 代理正在快速发展,导致出现多样化的评估基准、框架和协作协议。鉴于需要标准化评估和集成的日益增长的需求,我们系统性地将这些碎片化的努力整合为统一框架。然而,该领域仍然碎片化,缺乏统一的分类或全面调查。因此,我们对 2019 年至 2025 年期间开发的针对这些模型和代理的评估基准进行了横向比较,涵盖多个领域。此外,我们提出了约六十个基准的分类,涵盖通用和学术知识推理、数学问题解决、代码生成和软件工程、事实 grounding 和检索、特定领域评估、多模态和具身任务、任务编排和交互式评估。 Furthermore, we review AI-agent frameworks introduced between 2023 and 2025 that integrate large language models with modular toolkits to enable autonomous decision-making and multi-step reasoning. Moreover, we present real-world applications of autonomous AI agents in materials science, biomedical research, academic ideation, software engineering, synthetic data generation, chemical reasoning, mathematical problem-solving, geographic information systems, multimedia, healthcare, and finance. We then survey key agent-to-agent collaboration protocols, namely the Agent Communication Protocol (ACP), the Model Context Protocol (MCP), and the Agent-to-Agent Protocol (A2A). Finally, we discuss recommendations for future research, focusing on advanced reasoning strategies, failure modes in multi-agent LLM systems, automated scientific discovery, dynamic tool integration via reinforcement learning, integrated search capabilities, and security vulnerabilities in agent protocols.

关键词

引用

@article{arxiv.2504.19678,
  title  = {From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review},
  author = {Mohamed Amine Ferrag and Norbert Tihanyi and Merouane Debbah},
  journal= {arXiv preprint arXiv:2504.19678},
  year   = {2026}
}