中文
相关论文

相关论文: Measuring AI Ability to Complete Long Software Tas…

200 篇论文

We reframe the analysis of progress in AI by incorporating into an overall framework both the task performance of a system, and the time and resource costs incurred in the development and deployment of the system. These costs include: data,…

Artificial intelligence (AI) systems have become increasingly popular in many areas. Nevertheless, AI technologies are still in their developing stages, and many issues need to be addressed. Among those, the reliability of AI systems needs…

软件工程 · 计算机科学 2021-11-11 Yili Hong , Jiayi Lian , Li Xu , Jie Min , Yueyao Wang , Laura J. Freeman , Xinwei Deng

In this study, we explored the progression trajectories of artificial intelligence (AI) systems through the lens of complexity theory. We challenged the conventional linear and exponential projections of AI advancement toward Artificial…

A rising vision for AI in the open world centers on the development of systems that can complement humans for perceptual, diagnostic, and reasoning tasks. To date, systems aimed at complementing the skills of people have employed models…

人工智能 · 计算机科学 2020-05-05 Bryan Wilder , Eric Horvitz , Ece Kamar

Human Factors, Cognitive Engineering, and Human-Automation Interaction (HAI) form a trifecta, where users and technological systems of ever increasing autonomous control occupy a centre position. But with great autonomy comes great…

人机交互 · 计算机科学 2025-03-11 Gonçalo Hora de Carvalho

The efficiency of an AI system is contingent upon its ability to align with the specified requirements of a given task. How-ever, the inherent complexity of tasks often introduces the potential for harmful implications or adverse actions.…

计算机与社会 · 计算机科学 2023-12-08 Kamalakar Karlapalem

Large Language Models (LLMs) can generate code, but can they generate fast code for complex, real-world software systems? In this study, we investigate this question using a dataset of 65 tasks mined from performance-critical open-source…

软件工程 · 计算机科学 2026-04-10 Lirong Yi , Gregory Gay , Philipp Leitner

By defining the current limits (and thereby the frontiers), many boundaries are shaping, and will continue to shape, the future of Artificial Intelligence (AI). We push on these boundaries in order to make further progress into what were…

人工智能 · 计算机科学 2022-05-27 Ryan Watkins , Soheil Human

Large language models (LLMs) have the potential to boost human productivity by speeding up task completion -- provided users know when to offload cognitive work to them. But we do not know if users are well-calibrated in estimating these…

计算机与社会 · 计算机科学 2026-05-25 Sunny Yu , Myra Cheng , Ahmad Jabbar , Ilia Sucholutsky , Katherine M. Collins , Dan Jurafsky , Robert D. Hawkins

Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites. Since…

We present a quantitative model for tracking dangerous AI capabilities over time. Our goal is to help the policy and research community visualise how dangerous capability testing can give us an early warning about approaching AI risks. We…

人工智能 · 计算机科学 2024-12-23 Paolo Bova , Alessandro Di Stefano , The Anh Han

We interact with computers on an everyday basis, be it in everyday life or work, and many aspects of work can be done entirely with access to a computer and the Internet. At the same time, thanks to improvements in large language models…

Recent work has proposed artificial intelligence (AI) models that can learn to decide whether to make a prediction for an instance of a task or to delegate it to a human by considering both parties' capabilities. In simulations with…

人机交互 · 计算机科学 2023-03-17 Patrick Hemmer , Monika Westphal , Max Schemmer , Sebastian Vetter , Michael Vössing , Gerhard Satzger

In many real-world continuous action domains, human agents must decide which actions to attempt and then execute those actions to the best of their ability. However, humans cannot execute actions without error. Human performance in these…

人工智能 · 计算机科学 2024-08-21 Delma Nieves-Rivera , Christopher Archibald

Rapidly increasing AI capabilities have substantial real-world consequences, ranging from AI safety concerns to labor market consequences. The Model Evaluation & Threat Research (METR) report argues that AI capabilities have exhibited…

人工智能 · 计算机科学 2026-02-09 Haosen Ge , Hamsa Bastani , Osbert Bastani

Benchmarks are the primary tool for assessing progress in artificial intelligence (AI), yet current practice evaluates models on isolated test suites and provides little guidance for reasoning about generality or autonomous…

人工智能 · 计算机科学 2025-12-05 Przemyslaw Chojecki

People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learning, tracks progress, and prioritizes the other person's growth over immediate results. In contrast,…

人工智能 · 计算机科学 2026-04-08 Grace Liu , Brian Christian , Tsvetomira Dumbalska , Michiel A. Bakker , Rachit Dubey

AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the tendencies of models to exhibit particular behaviours - play a…

AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research. To spur the development of useful agents, we need benchmarks that are challenging, but more crucially, directly…

计算与语言 · 计算机科学 2024-09-18 Zachary S. Siegel , Sayash Kapoor , Nitya Nagdir , Benedikt Stroebl , Arvind Narayanan