English
Related papers

Related papers: Open-World Evaluations for Measuring Frontier AI C…

200 papers

With many organizations struggling to gain value from AI deployments, pressure to evaluate AI in an informed manner has intensified. Status quo AI evaluation approaches often mask the operational realities that ultimately determine…

Artificial Intelligence · Computer Science 2026-05-11 Matthew Holmes , Thiago Lacerda , Reva Schwartz

Large Language Models are commonly judged by their scores on standard benchmarks, yet such scores often overstate real capability since they mask the mix of skills a task actually demands. For example, ARC is assumed to test reasoning,…

Computation and Language · Computer Science 2025-10-03 Dongjun Kim , Gyuho Shim , Yongchan Chun , Minhyuk Kim , Chanjun Park , Heuiseok Lim

Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose a framework that calibrates items against the 'world…

Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Binary scores and static rubrics fail to capture these context-dependent requirements. Existing methods…

Computation and Language · Computer Science 2026-03-26 Shanghua Gao , Yuchang Su , Pengwei Sui , Curtis Ginder , Marinka Zitnik

Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and timeliness. Recent advances in AI have led the community to consider its use in peer…

Recent work proposes using world models to generate controlled virtual environments in which AI agents can be tested before deployment to ensure their reliability and safety. However, accurate world models often have high computational…

Artificial Intelligence · Computer Science 2025-04-08 Fernando Rosas , Alexander Boyd , Manuel Baltieri

For Large Language Models (LLMs), a disconnect persists between benchmark performance and real-world utility. Current evaluation frameworks remain fragmented, prioritizing technical metrics while neglecting holistic assessment for…

Artificial Intelligence · Computer Science 2025-11-19 Jun Wang , Ninglun Gu , Kailai Zhang , Zijiao Zhang , Yelun Bao , Jin Yang , Xu Yin , Liwei Liu , Yihuan Liu , Pengyong Li , Gary G. Yen , Junchi Yan

Cyber Ranges (CRs) have emerged as prominent platforms for cybersecurity training and education, especially for Critical Infrastructure (CI) sectors that face rising cyber threats. One way to address these threats is through hands-on…

Cryptography and Security · Computer Science 2025-12-12 Vyron Kampourakis , Georgios Kavallieratos , Georgios Spathoulas , Vasileios Gkioulos , Sokratis Katsikas

Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments,…

Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing agent benchmarks are limited by trajectory-opaque grading, underspecified safety and…

Artificial Intelligence · Computer Science 2026-05-08 Bowen Ye , Rang Li , Qibin Yang , Yuanxin Liu , Linli Yao , Hanglong Lv , Zhihui Xie , Chenxin An , Lei Li , Lingpeng Kong , Qi Liu , Zhifang Sui , Tong Yang

Large pre-trained language models have shown promise for few-shot learning, completing text-based tasks given only a few task-specific examples. Will models soon solve classification tasks that have so far been reserved for human research…

Frontier AI developers operate at the intersection of rapid technical progress, extreme risk exposure, and growing regulatory scrutiny. While a range of external evaluations and safety frameworks have emerged, comparatively little attention…

Computers and Society · Computer Science 2025-12-19 Francesca Gomez , Adam Buick , Leah Ferentinos , Haelee Kim , Elley Lee

LLM-based agents are increasingly expected to handle real-world assistant tasks, yet existing benchmarks typically evaluate them under isolated sources of difficulty, such as a single environment or fully specified instructions. This leaves…

Computation and Language · Computer Science 2026-04-16 Xiang Long , Li Du , Yilong Xu , Fangcheng Liu , Haoqing Wang , Ning Ding , Ziheng Li , Jianyuan Guo , Yehui Tang

Open-weight advanced AI models -- systems whose parameters are freely available for download and adaptation -- are reshaping the global AI landscape. As these models rapidly close the performance gap with closed alternatives, they enable…

Computers and Society · Computer Science 2026-02-24 Bengüsu Özcan , Alex Petropoulos , Max Reddel

Constructed-response questions are crucial to encourage generative processing and test a learner's understanding of core concepts. However, the limited availability of instructor time, large class sizes, and other resource constraints pose…

Computers and Society · Computer Science 2025-12-05 Shyam Agarwal , Ali Moghimi , Kevin C. Haudek

Automating AI research holds immense potential for accelerating scientific progress, yet current AI agents struggle with the complexities of rigorous, end-to-end experimentation. We introduce EXP-Bench, a novel benchmark designed to…

AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, and process-level approaches struggle to connect failure types to their precise locations…

This position paper argues that job exposure to AI should be measured with grounded, evidence-based methods, not inferred from LLM priors alone. Current theoretical exposure measures use zero-shot prompting to classify task-level AI…

Information Retrieval · Computer Science 2026-05-18 Luca Mouchel , Pierre Bouquet , Yossi Sheffi

AI measurement science has a wide variety of methodologies and measurements for comparing AI systems, resulting in what often appear to be "apples-to-oranges" comparisons across AI evaluations. To move toward "apples-to-apples" comparisons…

Human-Computer Interaction · Computer Science 2026-05-11 Yee-Yin Choong , Kristen Greene , Alice Qian , Meryem Marasli , Ziqi Yang , Sophia Chen , Laura Dabbish , Anand Rao , Hong Shen

Long-horizon, repetitive workflows are common in professional settings, such as processing expense reports from receipts and entering student grades from exam papers. These tasks are often tedious for humans since they can extend to extreme…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Jing Wu , Daphne Barretto , Yiye Chen , Nicholas Gydé , Yanan Jian , Yuhang He , Vibhav Vineet
‹ Prev 1 3 4 5 6 7 10 Next ›