English
Related papers

Related papers: GDPval: Evaluating AI Model Performance on Real-Wo…

200 papers

We study the staggered introduction of a generative AI-based conversational assistant using data from 5,172 customer support agents. Access to AI assistance increases worker productivity, as measured by issues resolved per hour, by 15\% on…

General Economics · Economics 2024-11-07 Erik Brynjolfsson , Danielle Li , Lindsey Raymond

Economic issues, such as inflation, energy costs, taxes, and interest rates, are a constant presence in our daily lives and have been exacerbated by global events such as pandemics, environmental disasters, and wars. A sustained history of…

Artificial Intelligence · Computer Science 2023-02-21 Abeer Abdullah Alaql , Fahad Alqurashi , Rashid Mehmood

The rapid advances in automation technologies, such as artificial intelligence (AI) and robotics, pose an increasing risk of automation for occupations, with a likely significant impact on the labour market. Recent social-economic studies…

Computers and Society · Computer Science 2022-09-07 Dawei Xu , Haoran Yang , Marian-Andrei Rizoiu , Guandong Xu

We study how Generative AI (GenAI) adoption is reshaping work. While prior studies show that GenAI enhances role-level productivity and task composition, its influence on skills - the fundamental enablers of task execution, and the ultimate…

General Economics · Economics 2025-06-17 Piyush Gulati , Arianna Marchetti , Phanish Puranam , Victoria Sevcenko

Compared to classical machine learning (ML) models, generative models offer a new usage paradigm where (i) a single model can be used for many different tasks out-of-the-box; (ii) users interact with this model over a series of natural…

Computer Science and Game Theory · Computer Science 2024-11-06 Rafid Mahmood

Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. However, existing benchmarks remain focused on single agentic capability, failing to capture…

Artificial Intelligence · Computer Science 2026-04-24 Keyu Li , Junhao Shi , Yang Xiao , Mohan Jiang , Jie Sun , Yunze Wu , Dayuan Fu , Shijie Xia , Xiaojie Cai , Tianze Xu , Weiye Si , Wenjie Li , Dequan Wang , Pengfei Liu

As AI adoption accelerates, research on its economic impacts becomes a salient source to consider for stakeholders of AI policy. Such research is however still in its infancy, and one in need of review. This paper aims to accomplish just…

General Economics · Economics 2024-12-09 Rafael Andersson Lipcsey

We formalize a macro-financial stress test for rapid AI adoption. Rather than a productivity bust or existential risk, we identify a distribution-and-contract mismatch: AI-generated abundance coexists with demand deficiency because economic…

Artificial Intelligence · Computer Science 2026-03-11 Xupeng Chen

Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing agent benchmarks are limited by trajectory-opaque grading, underspecified safety and…

Artificial Intelligence · Computer Science 2026-05-08 Bowen Ye , Rang Li , Qibin Yang , Yuanxin Liu , Linli Yao , Hanglong Lv , Zhihui Xie , Chenxin An , Lei Li , Lingpeng Kong , Qi Liu , Zhifang Sui , Tong Yang

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains, including software development, education, and technical assistance. Among these, software development is one of the key areas where LLMs are…

Computation and Language · Computer Science 2026-01-07 Inpyo Song , Eunji Jeon , Jangwon Lee

As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing…

Beyond scratch coding, exploiting large-scale code repositories (e.g., GitHub) for practical tasks is vital in real-world software development, yet current benchmarks rarely evaluate code agents in such authentic, workflow-driven scenarios.…

Given rapid progress toward advanced AI and risks from frontier AI systems (advanced AI systems pushing the boundaries of the AI capabilities frontier), the creation and implementation of AI governance and regulatory schemes deserves…

This study explores Large Language Models (LLMs) as autonomous agents for real-world tasks, including freelance software development. This work presents a new benchmark that evaluates LLMs on freelance programming and data analysis tasks…

Artificial Intelligence · Computer Science 2025-05-21 David Noever , Forrest McKee

As artificial intelligence (AI) systems approach and surpass expert human performance across a broad range of tasks, obtaining high-quality human supervision for evaluation and training becomes increasingly challenging. Our focus is on…

Machine Learning · Computer Science 2026-02-25 Ren Yin , Takashi Ishida , Masashi Sugiyama

We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step…

As Large Language Models reshape the global labor market, policymakers and workers need empirical data on which occupational skills may be most susceptible to automation. We present the Skill Automation Feasibility Index (SAFI),…

Computation and Language · Computer Science 2026-04-09 Rudra Jadhav , Janhavi Danve

Using the new data from the OECD-WTO world network of economic activities we construct the Google matrix $G$ of this directed network and perform its detailed analysis. The network contains 58 countries and 37 activity sectors for years…

Statistical Finance · Quantitative Finance 2015-07-21 V. Kandiah , H. Escaith , D. L. Shepelyansky

We present two comprehensive benchmarks to evaluate the performance of language models in coding assistance tasks, covering code writing, debugging, code review, and conceptual understanding. Our main contribution includes two curated…

Software Engineering · Computer Science 2024-12-10 Nidhish Shah , Zulkuf Genc , Dogu Araci

The accelerated evolution of large language models has raised questions about their comparative performance across domains of practical importance. GPT-4 by OpenAI introduced advances in reasoning, multimodality, and task generalization,…

Human-Computer Interaction · Computer Science 2025-08-28 Georgios P. Georgiou