中文
相关论文

相关论文: Remote Labor Index: Measuring AI Automation of Rem…

200 篇论文

AI agents are AI systems that can achieve complex goals autonomously. Assessing the level of agent autonomy is crucial for understanding both their potential benefits and risks. Current assessments of autonomy often focus on specific risks…

人工智能 · 计算机科学 2025-02-24 Peter Cihon , Merlin Stein , Gagan Bansal , Sam Manning , Kevin Xu

In recent years Artificial Intelligence (AI) has gained much popularity, with the scientific community as well as with the public. AI is often ascribed many positive impacts for different social domains such as medicine and the economy. On…

计算机与社会 · 计算机科学 2021-02-01 Kimon Kieslich , Marco Lünich , Frank Marcinkowski

CTI-REALM (Cyber Threat Real World Evaluation and LLM Benchmarking) is a benchmark designed to evaluate AI agents' ability to interpret cyber threat intelligence (CTI) and develop detection rules. The benchmark provides a realistic…

密码学与安全 · 计算机科学 2026-03-18 Arjun Chakraborty , Sandra Ho , Adam Cook , Manuel Meléndez

Amongst the most common use cases of modern AI is LLM chat with web search enabled. However, no direct evaluations of the quality of web research agents exist that control for the continually-changing web. We introduce Deep Research Bench,…

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon.…

Large Language Models (LLMs) show promise as data analysis agents, but existing benchmarks overlook the iterative nature of the field, where experts' decisions evolve with deeper insights of the dataset. To address this, we introduce…

计算与语言 · 计算机科学 2025-06-09 Hanyu Li , Haoyu Liu , Tingyu Zhu , Tianyu Guo , Zeyu Zheng , Xiaotie Deng , Michael I. Jordan

As industry reports claim agentic AI systems deliver double-digit productivity gains and multi-trillion dollar economic potential, the validity of these claims has become critical for investment decisions, regulatory policy, and responsible…

AI research agents accelerate ML research by automating hypothesis generation, experimentation, and empirical refinement. Existing agent strategies range from greedy hill-climbing to tree search and evolutionary optimization, yet which…

Comparing model performances on benchmark datasets is an integral part of measuring and driving progress in artificial intelligence. A model's performance on a benchmark dataset is commonly assessed based on a single or a small set of…

人工智能 · 计算机科学 2021-11-09 Kathrin Blagec , Georg Dorffner , Milad Moradi , Matthias Samwald

AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities. This raises a deeper question: can these systems extend their capabilities to automate AI…

Publicly accessible benchmarks that allow for assessing and comparing model performances are important drivers of progress in artificial intelligence (AI). While recent advances in AI capabilities hold the potential to transform medical…

人工智能 · 计算机科学 2022-12-26 Kathrin Blagec , Jakob Kraiger , Wolfgang Frühwirt , Matthias Samwald

Generative AI is altering work processes, task composition, and organizational design, yet its effects on employment and the macroeconomy remain unresolved. In this review, we synthesize theory and empirical evidence at three levels. First,…

综合经济学 · 经济学 2025-09-22 R. Maria del Rio-Chanona , Ekkehard Ernst , Rossana Merola , Daniel Samaan , Ole Teutloff

With generative AI emerging as a general-purpose technology, understanding its economic effects is among society's most pressing questions. Existing studies of AI impact have largely relied on predictions of AI capabilities or focused…

人工智能 · 计算机科学 2025-12-23 Kiran Tomlinson , Sonia Jaffe , Will Wang , Scott Counts , Siddharth Suri

AI-assisted research is crossing a threshold: fully automated systems can now generate research papers for as little as $15, while long-horizon agents can execute experiments, draft manuscripts, and simulate critique with minimal human…

Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These…

Contemporary benchmarks for agentic artificial intelligence (AI) frequently evaluate safety through isolated task-level accuracy thresholds, implicitly treating autonomous systems as single points of failure. This single-channel paradigm…

计算机与社会 · 计算机科学 2026-02-24 Nelu D. Radpour

LLM-based software engineering is influencing modern software development. In addition to correctness, prior studies have also examined the performance of software artifacts generated by AI agents. However, it is unclear how exactly the…

软件工程 · 计算机科学 2026-01-01 Md Nahidul Islam Opu , Shahidul Islam , Muhammad Asaduzzaman , Shaiful Chowdhury

This position paper argues that job exposure to AI should be measured with grounded, evidence-based methods, not inferred from LLM priors alone. Current theoretical exposure measures use zero-shot prompting to classify task-level AI…

信息检索 · 计算机科学 2026-05-18 Luca Mouchel , Pierre Bouquet , Yossi Sheffi

As artificial intelligence (AI) agents are deployed across economic domains, understanding their strategic behavior and market-level impact becomes critical. This paper puts forward a groundbreaking new framework that is the first to…

多智能体系统 · 计算机科学 2025-12-05 Christopher Chiu , Simpson Zhang , Mihaela van der Schaar

Deep reinforcement learning (RL) algorithms are predominantly evaluated by comparing their relative performance on a large suite of tasks. Most published results on deep RL benchmarks compare point estimates of aggregate performance such as…

机器学习 · 计算机科学 2022-01-06 Rishabh Agarwal , Max Schwarzer , Pablo Samuel Castro , Aaron Courville , Marc G. Bellemare