中文
相关论文

相关论文: Nonstandard Errors in AI Agents

200 篇论文

Forecasting when AI systems will become capable of meaningfully accelerating AI research is a central challenge for AI safety. Existing benchmarks measure broad capability growth, but may not provide ample early warning signals for…

多智能体系统 · 计算机科学 2026-04-30 Joshua Sherwood , Ben Aybar , Benjamin Kaplan

As autonomous AI agents are deployed in persistent, interacting networks -- coordinating tasks, routing resources, and accumulating reputational histories -- the social dynamics that emerge will determine who receives opportunity and who…

人工智能 · 计算机科学 2026-05-28 Messi H. J. Lee

Distributed AI inference pipelines rely heavily on timestamp-based observability to understand system behavior. This work demonstrates that even small clock skew between nodes can cause observability to become causally incorrect while the…

人工智能 · 计算机科学 2026-04-24 Ankur Sharma , Deep Shah , David Lariviere , Hesham ElBakoury

Aligning AI systems with human values remains a fundamental challenge, but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? We study a strategic setting where a human user interacts with…

机器学习 · 计算机科学 2026-02-04 Natalie Collina , Surbhi Goel , Aaron Roth , Emily Ryu , Mirah Shi

Background: Large language models are typically evaluated as models, benchmarks, or short conversational episodes. Less is known about what happens when an agent is embedded persistently in a real academic research environment with durable…

多智能体系统 · 计算机科学 2026-05-27 Anas H. Alzahrani

An important goal in the field of human-AI interaction is to help users more appropriately trust AI systems' decisions. A situation in which the user may particularly benefit from more appropriate trust is when the AI receives anomalous…

人机交互 · 计算机科学 2022-04-29 Marissa Radensky , Dustin Burson , Rajya Bhaiya , Daniel S. Weld

How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article…

计算机与社会 · 计算机科学 2026-05-13 Jason Miklian , Kristian Hoelscher , John E. Katsos

AI agents are increasingly being deployed to automate tasks, often based on underspecified user instructions. Making unwarranted assumptions to compensate for the missing information and failing to ask clarifying questions can lead to…

人工智能 · 计算机科学 2026-02-24 Sanidhya Vijayvargiya , Xuhui Zhou , Akhila Yerukola , Maarten Sap , Graham Neubig

AI coding agents demonstrate strong performance on general-purpose software benchmarks. However, their ability to handle 5G network engineering tasks remains unexplored. We propose SWE-Bench~5G, the first benchmark designed to investigate…

网络与互联网体系结构 · 计算机科学 2026-04-30 Jiao Chen , Jianhua Tang , Xiaotong Yang , Zuohong Lv

According to canonical negotiation theory, people's success in a negotiation depends on how well they balance competing demands--empathizing and asserting, demonstrating concern for other and concern for self, being soft on the people and…

人工智能 · 计算机科学 2026-05-21 Michelle A. Vaccaro , Jared R. Curhan

The rapid adoption of AI agents across domains has made systematic evaluation crucial for ensuring their usefulness and successful production deployment. Evaluation of AI agents typically involves using a fixed set of benchmarks and…

Agents, language model-based systems capable of reasoning, planning, and acting are widely adopted in real-world tasks, yet how their performance changes as these systems scale across key dimensions remains underexplored. We introduce…

Ambient AI "scribe" systems promise to reduce clinical documentation burden, but automatic speech recognition (ASR) errors can remain unnoticed without careful review, and high-quality human reference transcripts are often unavailable for…

声音 · 计算机科学 2026-04-17 Abdolamir Karbalaie , Fernando Seoane , Farhad Abtahi

Enterprise AI systems, built on large language models, retrieval pipelines and autonomous agents, introduce a class of risks that traditional software quality assurance was never designed to address. These systems are probabilistic,…

软件工程 · 计算机科学 2026-05-25 Chitra Badagi , Divye Singh , Animesh Sen , Adinath Shirsath

This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries and fields. Using large-scale data on journal publications from Dimensions, we construct…

计算与语言 · 计算机科学 2026-05-27 Shang Wu , Randol Yao

Computational social science lacks a scalable and reliable mechanism to assure quality for AI-assisted qualitative coding when tasks demand domain expertise and long-text reasoning, and traditional double-coding is prohibitively costly at…

计算机与社会 · 计算机科学 2025-10-01 Zhilong Zhao , Yindi Liu

AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, and process-level approaches struggle to connect failure types to their precise locations…

Auditing plays a pivotal role in the development of trustworthy AI. However, current research primarily focuses on creating auditable AI documentation, which is intended for regulators and experts rather than end-users affected by AI…

计算机与社会 · 计算机科学 2023-05-31 Nicolas Scharowski , Michaela Benk , Swen J. Kühne , Léane Wettstein , Florian Brühlmann

A major bottleneck in characterizing the failure modes of generative AI systems is the cost and time of annotation and evaluation. Consequently, adaptive testing paradigms have gained popularity, where one opportunistically decides which…

人工智能 · 计算机科学 2026-05-11 Siyu Zhou , Patrick Vossler , Venkatesh Sivaraman , Yifan Mai , Jean Feng

Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. However, existing benchmarks remain focused on single agentic capability, failing to capture…

‹ 上一页 1 8 9 10 下一页 ›