Artificial Intelligence · Computer Science
Survey on Evaluation of LLM-based Agents
Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel +4
2026-04-24
Software Engineering · Computer Science
Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye +1
2025-11-07
Software Engineering · Computer Science
LLM-Agent-UMF: LLM-based Agent Unified Modeling Framework for Seamless Design of Multi Active/Passive Core-Agent Architectures
Amine Ben Hassouna, Hana Chaari, Ines Belhaj
2025-11-24
Machine Learning · Computer Science
Evaluation and Benchmarking of LLM Agents: A Survey
Mahmoud Mohammadi, Yipeng Li, Jane Lo, Wendy Yip
2025-07-30
Artificial Intelligence · Computer Science
Fundamentals of Building Autonomous LLM Agents
Victor de Lamo Castrillo, Habtom Kahsay Gidey, Alexander Lenz, Alois Knoll
2025-10-13
Machine Learning · Computer Science
Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring
Gusseppe Bravo-Rocca, Peini Liu, Jordi Guitart, Rodrigo M Carrillo-Larco +2
2025-06-12
Computer Vision and Pattern Recognition · Computer Science
Building LLM Agents by Incorporating Insights from Computer Systems
Yapeng Mi, Zhi Gao, Xiaojian Ma, Qing Li
2025-04-08
Computation and Language · Computer Science
Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
Yihong Tang, Kehai Chen, Liang Yue, Jinxin Fan +10
2025-10-21
Computation and Language · Computer Science
Large Language Model Agent: A Survey on Methodology, Applications and Challenges
Junyu Luo, Weizhi Zhang, Ye Yuan, Yusheng Zhao +22
2025-03-28
Artificial Intelligence · Computer Science
Towards a Science of Scaling Agent Systems
Yubin Kim, Ken Gu, Chanwoo Park, Chunjong Park +16
2026-04-10
Software Engineering · Computer Science
Automated structural testing of LLM-based agents: methods, framework, and case studies
Jens Kohl, Otto Kruse, Youssef Mostafa, Andre Luckow +8
2026-01-28
Artificial Intelligence · Computer Science
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
Chris Ge, Daria Kryvosheieva, Daniel Fried, Uzay Girit +1
2026-04-02
Artificial Intelligence · Computer Science
AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
Luca Gioacchini, Giuseppe Siracusano, Davide Sanvito, Kiril Gashteovski +3
2024-04-10
Multiagent Systems · Computer Science
LLM-Enabled Multi-Agent Systems: Empirical Evaluation and Insights into Emerging Design Patterns & Paradigms
Harri Renney, Maxim N Nethercott, Nathan Renney, Peter Hayes
2026-01-08
Software Engineering · Computer Science
From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
Haolin Jin, Linghan Huang, Haipeng Cai, Jun Yan +2
2025-04-15
Software Engineering · Computer Science
Supporting architecture evaluation for ATAM scenarios with LLMs
Rafael Capilla, J. Andrés Díaz-Pace, Yamid Ramírez, Jennifer Pérez +1
2025-06-03
Computation and Language · Computer Science
Investigating Agency of LLMs in Human-AI Collaboration Tasks
Ashish Sharma, Sudha Rao, Chris Brockett, Akanksha Malhotra +2
2024-02-09
Artificial Intelligence · Computer Science
From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing
Lanxiao Huang, Daksh Dave, Tyler Cody, Peter Beling +1
2025-11-14