中文
相关论文

相关论文: GDPval: Evaluating AI Model Performance on Real-Wo…

200 篇论文

We study the staggered introduction of a generative AI-based conversational assistant using data from 5,172 customer support agents. Access to AI assistance increases worker productivity, as measured by issues resolved per hour, by 15\% on…

综合经济学 · 经济学 2024-11-07 Erik Brynjolfsson , Danielle Li , Lindsey Raymond

Economic issues, such as inflation, energy costs, taxes, and interest rates, are a constant presence in our daily lives and have been exacerbated by global events such as pandemics, environmental disasters, and wars. A sustained history of…

人工智能 · 计算机科学 2023-02-21 Abeer Abdullah Alaql , Fahad Alqurashi , Rashid Mehmood

The rapid advances in automation technologies, such as artificial intelligence (AI) and robotics, pose an increasing risk of automation for occupations, with a likely significant impact on the labour market. Recent social-economic studies…

计算机与社会 · 计算机科学 2022-09-07 Dawei Xu , Haoran Yang , Marian-Andrei Rizoiu , Guandong Xu

We study how Generative AI (GenAI) adoption is reshaping work. While prior studies show that GenAI enhances role-level productivity and task composition, its influence on skills - the fundamental enablers of task execution, and the ultimate…

综合经济学 · 经济学 2025-06-17 Piyush Gulati , Arianna Marchetti , Phanish Puranam , Victoria Sevcenko

Compared to classical machine learning (ML) models, generative models offer a new usage paradigm where (i) a single model can be used for many different tasks out-of-the-box; (ii) users interact with this model over a series of natural…

计算机科学与博弈论 · 计算机科学 2024-11-06 Rafid Mahmood

Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. However, existing benchmarks remain focused on single agentic capability, failing to capture…

As AI adoption accelerates, research on its economic impacts becomes a salient source to consider for stakeholders of AI policy. Such research is however still in its infancy, and one in need of review. This paper aims to accomplish just…

综合经济学 · 经济学 2024-12-09 Rafael Andersson Lipcsey

We formalize a macro-financial stress test for rapid AI adoption. Rather than a productivity bust or existential risk, we identify a distribution-and-contract mismatch: AI-generated abundance coexists with demand deficiency because economic…

人工智能 · 计算机科学 2026-03-11 Xupeng Chen

Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing agent benchmarks are limited by trajectory-opaque grading, underspecified safety and…

人工智能 · 计算机科学 2026-05-08 Bowen Ye , Rang Li , Qibin Yang , Yuanxin Liu , Linli Yao , Hanglong Lv , Zhihui Xie , Chenxin An , Lei Li , Lingpeng Kong , Qi Liu , Zhifang Sui , Tong Yang

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains, including software development, education, and technical assistance. Among these, software development is one of the key areas where LLMs are…

计算与语言 · 计算机科学 2026-01-07 Inpyo Song , Eunji Jeon , Jangwon Lee

As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing…

Beyond scratch coding, exploiting large-scale code repositories (e.g., GitHub) for practical tasks is vital in real-world software development, yet current benchmarks rarely evaluate code agents in such authentic, workflow-driven scenarios.…

Given rapid progress toward advanced AI and risks from frontier AI systems (advanced AI systems pushing the boundaries of the AI capabilities frontier), the creation and implementation of AI governance and regulatory schemes deserves…

This study explores Large Language Models (LLMs) as autonomous agents for real-world tasks, including freelance software development. This work presents a new benchmark that evaluates LLMs on freelance programming and data analysis tasks…

人工智能 · 计算机科学 2025-05-21 David Noever , Forrest McKee

As artificial intelligence (AI) systems approach and surpass expert human performance across a broad range of tasks, obtaining high-quality human supervision for evaluation and training becomes increasingly challenging. Our focus is on…

机器学习 · 计算机科学 2026-02-25 Ren Yin , Takashi Ishida , Masashi Sugiyama

We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step…

As Large Language Models reshape the global labor market, policymakers and workers need empirical data on which occupational skills may be most susceptible to automation. We present the Skill Automation Feasibility Index (SAFI),…

计算与语言 · 计算机科学 2026-04-09 Rudra Jadhav , Janhavi Danve

Using the new data from the OECD-WTO world network of economic activities we construct the Google matrix $G$ of this directed network and perform its detailed analysis. The network contains 58 countries and 37 activity sectors for years…

统计金融 · 定量金融 2015-07-21 V. Kandiah , H. Escaith , D. L. Shepelyansky

We present two comprehensive benchmarks to evaluate the performance of language models in coding assistance tasks, covering code writing, debugging, code review, and conceptual understanding. Our main contribution includes two curated…

软件工程 · 计算机科学 2024-12-10 Nidhish Shah , Zulkuf Genc , Dogu Araci

The accelerated evolution of large language models has raised questions about their comparative performance across domains of practical importance. GPT-4 by OpenAI introduced advances in reasoning, multimodality, and task generalization,…

人机交互 · 计算机科学 2025-08-28 Georgios P. Georgiou