English
Related papers

Related papers: SEA-Eval: A Benchmark for Evaluating Self-Evolving…

200 papers

Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models. However, existing methods rely solely on outcome rewards, without explicitly optimizing verification or leveraging…

Software Engineering · Computer Science 2025-10-22 Yiyang Jin , Kunzhao Xu , Hang Li , Xueting Han , Yanmin Zhou , Cheng Li , Jing Bai

With the advancement of Agentic AI, researchers are increasingly leveraging autonomous agents to address challenges in software engineering (SE). However, the large language models (LLMs) that underpin these agents often function as black…

Software Engineering · Computer Science 2026-04-03 Jingyue Li , André Storhaug

Large Language Model (LLM)-based agents have achieved notable success on short-horizon and highly structured tasks. However, their ability to maintain coherent decision-making over long horizons in realistic and dynamic environments remains…

Artificial Intelligence · Computer Science 2026-03-18 Linghua Zhang , Jun Wang , Jingtong Wu , Zhisong Zhang

Interactive tool-using agents must solve real-world tasks via multi-turn interaction with both humans and external environments, requiring dialogue state tracking, multi-step tool execution, while following complex instructions.…

Artificial Intelligence · Computer Science 2026-03-11 Jiaxuan Gao , Jiaao Chen , Chuyi He , Shusheng Xu , Di Jin , Yi Wu

Traditional self-adaptive systems automatically reconfigure existing components in response to changing requirements, but provide limited support for the generation of novel functionalities. The software generation capabilities of large…

Software Engineering · Computer Science 2026-04-21 Md Asif Iqbal Fahim , Oluwadamilola Adebayo , Alessio Ferrari

Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot resolve. Rather than constructing harder alternatives, we ask whether existing tasks can…

Computation and Language · Computer Science 2026-05-29 Jiamin Chen , Yidi Wu , Qiexiang Wang , Qianben Chen , Yuchen Li , Yansen Zhang , Xiaokun Zhang , Wangchunshu Zhou , Chen Ma

Background: The surge in single-cell omics data exposes limitations in traditional, manually defined analysis workflows. AI agents offer a paradigm shift, enabling adaptive planning, executable code generation, traceable decisions, and…

Genomics · Quantitative Biology 2026-03-17 Yang Liu , Lu Zhou , Xiawei Du , Ruikun He , Xuguang Zhang , Rongbo Shen , Yixue Li

Most agents today ``self-evolve'' by following rewards and rules defined by humans. However, this process remains fundamentally dependent on external supervision; without human guidance, the evolution stops. In this work, we train agents to…

Artificial Intelligence · Computer Science 2026-04-21 Qifan Zhang , Dongyang Ma , Tianqing Fang , Jia Li , Jing Tang , Nuo Chen , Haitao Mi , Yan Wang

Action chunking has recently emerged as a standard practice in flow-based Vision-Language-Action (VLA) models. However, the effect and choice of the execution horizon - the number of actions to be executed from each predicted chunk -…

Robotics · Computer Science 2026-02-26 Haoxuan Wang , Gengyu Zhang , Yan Yan , Ramana Rao Kompella , Gaowen Liu

Due to the dynamically evolving nature of real-world query streams, relevance models struggle to generalize to practical search scenarios. A sophisticated solution is self-evolution techniques. However, in large-scale industrial settings…

Computation and Language · Computer Science 2026-04-21 Chenglong Wang , Canjia Li , Xingzhao Zhu , Yifu Huo , Huiyu Wang , Weixiong Lin , Yun Yang , Qiaozhi He , Tianhua Zhou , Xiaojia Chang , Jingbo Zhu , Tong Xiao

Large language models (LLMs) are increasingly deployed as customer-facing agents, yet evaluating their reliability remains challenging due to stochastic, multi-turn interactions. Current evaluation protocols rely on linear Monte Carlo…

Artificial Intelligence · Computer Science 2026-04-24 Itay Nakash , George Kour , Ateret Anaby-Tavor

Evaluating AI agents within complex, interactive environments that mirror real-world challenges is critical for understanding their practical capabilities. While existing agent benchmarks effectively assess skills like tool use or…

Artificial Intelligence · Computer Science 2025-08-15 Long Phan , Mantas Mazeika , Andy Zou , Dan Hendrycks

AI Agents are changing the way work gets done, both in consumer and enterprise domains. However, the design patterns and architectures to build highly capable agents or multi-agent systems are still developing, and the understanding of the…

Artificial Intelligence · Computer Science 2024-07-19 Tamer Abuelsaad , Deepak Akkil , Prasenjit Dey , Ashish Jagmohan , Aditya Vempaty , Ravi Kokku

Autonomous agents can adapt their behaviour to changing environments, but remain bound to requirements, goals, and capabilities fixed at design time, preventing genuine software evolution. This paper introduces self-evolving software…

Software Engineering · Computer Science 2026-05-01 Marco Robol , Paolo Giorgini

Southeast Asia (SEA) is a region of extraordinary linguistic and cultural diversity, yet it remains significantly underrepresented in vision-language (VL) research. This often results in artificial intelligence (AI) models that fail to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Samuel Cahyawijaya , Holy Lovenia , Joel Ruben Antony Moniz , Tack Hwa Wong , Mohammad Rifqi Farhansyah , Thant Thiri Maung , Frederikus Hudi , David Anugraha , Muhammad Ravi Shulthan Habibi , Muhammad Reza Qorib , Amit Agarwal , Joseph Marvin Imperial , Hitesh Laxmichand Patel , Vicky Feliren , Bahrul Ilmi Nasution , Manuel Antonio Rufino , Genta Indra Winata , Rian Adam Rajagede , Carlos Rafael Catalan , Mohamed Fazli Imam , Priyaranjan Pattnayak , Salsabila Zahirah Pranida , Kevin Pratama , Yeshil Bangera , Adisai Na-Thalang , Patricia Nicole Monderin , Yueqi Song , Christian Simon , Lynnette Hui Xian Ng , Richardy Lobo' Sapan , Taki Hasan Rafi , Bin Wang , Supryadi , Kanyakorn Veerakanjana , Piyalitt Ittichaiwong , Matthew Theodore Roque , Karissa Vincentio , Takdanai Kreangphet , Phakphum Artkaew , Kadek Hendrawan Palgunadi , Yanzhi Yu , Rochana Prih Hastuti , William Nixon , Mithil Bangera , Adrian Xuan Wei Lim , Aye Hninn Khine , Hanif Muhammad Zhafran , Teddy Ferdinan , Audra Aurora Izzani , Ayushman Singh , Evan , Jauza Akbar Krito , Michael Anugraha , Fenal Ashokbhai Ilasariya , Haochen Li , John Amadeo Daniswara , Filbert Aurelian Tjiaranata , Eryawan Presma Yulianrifat , Can Udomcharoenchaikit , Fadil Risdian Ansori , Mahardika Krisna Ihsani , Giang Nguyen , Anab Maulana Barik , Dan John Velasco , Rifo Ahmad Genadi , Saptarshi Saha , Chengwei Wei , Isaiah Flores , Kenneth Ko Han Chen , Anjela Gail Santos , Wan Shen Lim , Kaung Si Phyo , Tim Santos , Meisyarah Dwiastuti , Jiayun Luo , Jan Christian Blaise Cruz , Ming Shan Hee , Ikhlasul Akmal Hanif , M. Alif Al Hakim , Muhammad Rizky Sya'ban , Kun Kerdthaisong , Lester James V. Miranda , Fajri Koto , Tirana Noor Fatyanosa , Alham Fikri Aji , Jostin Jerico Rosal , Jun Kevin , Robert Wijaya , Onno P. Kampman , Ruochen Zhang , Börje F. Karlsson , Peerat Limkonchotiwat

Despite the remarkable success of large language models (LLMs), they still face bottlenecks while deploying in dynamic, real-world settings with primary challenges being concept drift and the high cost of gradient-based adaptation.…

Artificial Intelligence · Computer Science 2026-05-21 Nitin Vetcha , Dianbo Liu

Evaluating large language model (LLM)-based multi-agent systems remains a critical challenge, as these systems must exhibit reliable coordination, transparent decision-making, and verifiable performance across evolving tasks. Existing…

Artificial Intelligence · Computer Science 2026-01-21 YenTing Lee , Keerthi Koneru , Zahra Moslemi , Sheethal Kumar , Ramesh Radhakrishnan

While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reasoning where errors are often rectifiable via backtracking, tool-use failures frequently induce…

Artificial Intelligence · Computer Science 2026-03-17 Shengda Fan , Xuyan Ye , Yupeng Huo , Zhi-Yuan Chen , Yiju Guo , Shenzhi Yang , Wenkai Yang , Shuqi Ye , Jingwen Chen , Haotian Chen , Xin Cong , Yankai Lin

Building generalist agents that can handle diverse tasks and evolve themselves across different environments is a long-term goal in the AI community. Large language models (LLMs) are considered a promising foundation to build such agents…

Much of the research on learning symbolic models of AI agents focuses on agents with stationary models. This assumption fails to hold in settings where the agent's capabilities may change as a result of learning, adaptation, or other…

Artificial Intelligence · Computer Science 2022-05-20 Rashmeet Kaur Nayyar , Pulkit Verma , Siddharth Srivastava