中文
相关论文

相关论文: Open-World Evaluations for Measuring Frontier AI C…

200 篇论文

The ability of Large Language Models (LLMs) to use external tools unlocks powerful real-world interactions, making rigorous evaluation essential. However, current benchmarks primarily report final accuracy, revealing what models can do but…

计算与语言 · 计算机科学 2026-01-29 Qihao Wang , Yue Hu , Mingzhe Lu , Jiayue Wu , Yanbing Liu , Yuanmin Tang

We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management…

The impact of frontier AI (i.e., AI agents and foundation models) in cybersecurity is rapidly increasing. In this paper, we comprehensively analyze this trend through multiple aspects: quantitative benchmarks, qualitative literature review,…

密码学与安全 · 计算机科学 2025-12-01 Yujin Potter , Wenbo Guo , Zhun Wang , Tianneng Shi , Hongwei Li , Andy Zhang , Patrick Gage Kelley , Kurt Thomas , Dawn Song

Quantitative Artificial Intelligence (AI) Benchmarks have emerged as fundamental tools for evaluating the performance, capability, and safety of AI models and systems. Currently, they shape the direction of AI development and are playing an…

The opacity of AI models necessitates both validation and evaluation before their integration into services. To investigate these models, explainable AI (XAI) employs methods that elucidate the relationship between input features and output…

密码学与安全 · 计算机科学 2024-10-02 Zerui Wang , Yan Liu

During the past decades, numerous successes of AI has been made on "specific capabilities", named closed-world, such as artificial environments or specific real-world tasks. This well-defined narrow capability brings two nice benefits, a…

机器学习 · 计算机科学 2025-07-08 Jianyu Zhang

Frontier artificial intelligence (AI) systems could pose increasing risks to public safety and security. But what level of risk is acceptable? One increasingly popular approach is to define capability thresholds, which describe AI…

计算机与社会 · 计算机科学 2024-06-24 Leonie Koessler , Jonas Schuett , Markus Anderljung

Benchmarking has long served as a foundational practice in machine learning and, increasingly, in modern AI systems such as large language models, where shared tasks, metrics, and leaderboards offer a common basis for measuring progress and…

人工智能 · 计算机科学 2026-02-16 Philip Waggoner

Current frontier AI safety evaluations emphasize static benchmarks, third-party annotations, and red-teaming. In this position paper, we argue that AI safety research should focus on human-centered evaluations that measure harmful…

计算机与社会 · 计算机科学 2026-03-31 Michelle Vaccaro , Jaeyoon Song , Abdullah Almaatouq , Michiel A. Bakker

Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, handwritten single-variable calculus work from UC Irvine. Using…

机器学习 · 计算机科学 2026-03-03 Zhiqi Yu , Xingping Liu , Haobin Mao , Mingshuo Liu , Long Chen , Jack Xin , Yifeng Yu

Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Yet skepticism about their reliability continues to grow. How can we know that a reported…

人工智能 · 计算机科学 2026-05-19 Nathanael Jo , Ashia Wilson

As AI models tackle increasingly complex problems, ensuring reliable human oversight becomes more challenging due to the difficulty of verifying solutions. Approaches to scaling AI supervision include debate, in which two agents engage in…

人工智能 · 计算机科学 2025-04-01 Gabriel Recchia , Chatrik Singh Mangat , Issac Li , Gayatri Krishnakumar

Artificial Intelligence (AI) Safety Institutes and governments worldwide are deciding whether they evaluate advanced AI themselves, support a private evaluation ecosystem or do both. Evaluation regimes have been established in a wide range…

计算机与社会 · 计算机科学 2025-08-06 Merlin Stein , Milan Gandhi , Theresa Kriecherbauer , Amin Oueslati , Robert Trager

To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons…

人工智能 · 计算机科学 2019-11-26 François Chollet

Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These…

Over the past year, artificial intelligence (AI) companies have been increasingly adopting AI safety frameworks. These frameworks outline how companies intend to keep the potential risks associated with developing and deploying frontier AI…

计算机与社会 · 计算机科学 2024-09-16 Jide Alaga , Jonas Schuett , Markus Anderljung

Research on new optimization algorithms is often funded based on the motivation that such algorithms might improve the capabilities to deal with real-world and industrially relevant optimization challenges. Besides a huge variety of…

神经与进化计算 · 计算机科学 2020-07-02 Ramses Sala , Ralf Müller

Benchmarks such as ARC, Raven-inspired tests, and the Blackbird Task are widely used to evaluate the intelligence of large language models (LLMs). Yet, the concept of intelligence remains elusive- lacking a stable definition and failing to…

人工智能 · 计算机科学 2025-11-18 Ruchira Dhar , Ninell Oldenburg , Anders Soegaard

There is a tendency across different subfields in AI to valorize a small collection of influential benchmarks. These benchmarks operate as stand-ins for a range of anointed common problems that are frequently framed as foundational…

机器学习 · 计算机科学 2021-12-01 Inioluwa Deborah Raji , Emily M. Bender , Amandalynne Paullada , Emily Denton , Alex Hanna