English
Related papers

Related papers: Re-evaluating Evaluation

200 papers

Current evaluations of agents remain centered around one-shot task completion, failing to account for the inherently iterative and collaborative nature of many real-world problems, where human goals are often underspecified and evolve. We…

This paper investigates the equilibrium convergence properties of a proposed algorithm for potential games with continuous strategy spaces in the presence of feedback delays, a main challenge in multi-agent systems that compromises the…

Optimization and Control · Mathematics 2023-03-20 Yuanhanqing Huang , Jianghai Hu

AI agents are increasingly deployed in production, yet their security evaluations remain bottlenecked by manual red-teaming or static benchmarks that fail to model adaptive, multi-turn adversaries. We propose NAAMSE, an evolutionary…

Artificial Intelligence · Computer Science 2026-03-10 Kunal Pai , Parth Shah , Harshil Patel

Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their capabilities. Current evaluation methodology, mostly based…

Artificial Intelligence · Computer Science 2024-01-01 Xiting Wang , Liming Jiang , Jose Hernandez-Orallo , David Stillwell , Luning Sun , Fang Luo , Xing Xie

We study collaborative normal mean estimation, where $m$ strategic agents collect i.i.d samples from a normal distribution $\mathcal{N}(\mu, \sigma^2)$ at a cost. They all wish to estimate the mean $\mu$. By sharing data with each other,…

Computer Science and Game Theory · Computer Science 2023-11-22 Yiding Chen , Xiaojin Zhu , Kirthevasan Kandasamy

Nash equilibrium is often heralded as a guiding principle for rational decision-making in strategic interactions. However, it is well-known that Nash equilibrium sometimes fails as a reliable predictor of outcomes, with two of the most…

Computer Science and Game Theory · Computer Science 2023-12-27 Ivan Geffner , Moshe Tennenholtz

The distribution of efficient individuals in the economy and the efforts that they will put in if they are hired, there are two important concerns for a technologically advanced firm. wants to open a new branch. The firm does not have…

Computer Science and Game Theory · Computer Science 2025-01-27 Sujata Goala , Mridu Prabal Goswami , Surajit Borkotokey

Here, we develop a deep learning algorithm for solving Principal-Agent (PA) mean field games with market-clearing conditions -- a class of problems that have thus far not been studied and one that poses difficulties for standard numerical…

Machine Learning · Computer Science 2021-10-05 Steven Campbell , Yichao Chen , Arvind Shrivats , Sebastian Jaimungal

We consider a repeatedly played generalized Nash equilibrium game. This induces a multi-agent online learning problem with joint constraints. An important challenge in this setting is that the feasible set for each agent depends on the…

Machine Learning · Computer Science 2024-10-04 Sarah Sachs , Hedi Hadiji , Tim van Erven , Mathias Staudigl

The software engineering research community faces a systemic crisis: peer review is failing under growing submissions, misaligned incentives, and reviewer fatigue. Community surveys reveal that researchers perceive the process as "broken."…

Multiagent Systems · Computer Science 2026-01-28 Ahmad Farooq , Kamran Iqbal

Neural architecture search has proven to be a powerful approach to designing and refining neural networks, often boosting their performance and efficiency over manually-designed variations, but comes with computational overhead. While there…

Continuous-time gradient-based Nash equilibrium seeking algorithms enjoy a passivity property under a suitable monotonicity assumption. This feature has been exploited to design distributed algorithms that converge to Nash equilibria and…

Optimization and Control · Mathematics 2018-12-04 Claudio De Persis , Sergio Grammatico

We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through \emph{latent world recovery}. AvalancheBench improves on existing benchmarks in three ways. First, it evaluates analytical understanding rather than…

The release of tabular benchmarks, such as NAS-Bench-101 and NAS-Bench-201, has significantly lowered the computational overhead for conducting scientific research in neural architecture search (NAS). Although they have been widely adopted…

AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in practice. This discrepancy highlights a fundamental limitation…

Artificial Intelligence · Computer Science 2026-02-24 Stephan Rabanser , Sayash Kapoor , Peter Kirgis , Kangheng Liu , Saiteja Utpala , Arvind Narayanan

We formulate a general framework for competitive gradient-based learning that encompasses a wide breadth of multi-agent learning algorithms, and analyze the limiting behavior of competitive gradient-based learning algorithms using dynamical…

Machine Learning · Computer Science 2020-02-21 Eric Mazumdar , Lillian J. Ratliff , S. Shankar Sastry

AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions of inquiry; indeed, there are now many such agents, ranging…

Averaging checkpoints along the training trajectory is a simple yet powerful approach to improve the generalization performance of Machine Learning models and reduce training time. Motivated by these potential gains, and in an effort to…

Machine Learning · Computer Science 2025-11-25 Niccolò Ajroldi , Antonio Orvieto , Jonas Geiping

The computational characterization of game-theoretic solution concepts is a central topic in artificial intelligence, with the aim of developing computationally efficient tools for finding optimal ways to behave in strategic interactions.…

Computer Science and Game Theory · Computer Science 2013-04-05 Nicola Gatti , Marco Rocco , Tuomas Sandholm

We consider multi-agent decision making, where each agent optimizes its cost function subject to constraints. Agents' actions belong to a compact convex Euclidean space and the agents' cost functions are coupled. We propose a distributed…

Optimization and Control · Mathematics 2016-12-01 Tatiana Tatarenko , Maryam Kamgarpour