中文
相关论文

相关论文: The Missing Red Line: How Commercial Pressure Erod…

200 篇论文

Large language models are being deployed as mental health support agents at scale, yet only 16% of LLM-based chatbot interventions have undergone rigorous clinical efficacy testing, and simulations reveal psychological deterioration in over…

计算与语言 · 计算机科学 2026-04-28 Suhas BN , Andrew M. Sherrill , Rosa I. Arriaga , Chris W. Wiese , Saeed Abdullah

Autonomous agents acting in the real-world often operate based on models that ignore certain aspects of the environment. The incompleteness of any given model -- handcrafted or machine acquired -- is inevitable due to practical limitations…

计算机与社会 · 计算机科学 2021-10-20 Sandhya Saisubramanian , Shlomo Zilberstein , Ece Kamar

While it is still unclear if agents with Artificial General Intelligence (AGI) could ever be built, we can already use mathematical models to investigate potential safety systems for these agents. We present an AGI safety layer that creates…

人工智能 · 计算机科学 2020-07-13 Koen Holtman

AI-based writing assistants are ubiquitous, yet little is known about how users' mental models shape their use. We examine two types of mental models -- functional or related to what the system does, and structural or related to how the…

人机交互 · 计算机科学 2026-04-08 Shalaleh Rismani , Su Lin Blodgett , Q. Vera Liao , Alexandra Olteanu , AJung Moon

Large language models deployed as agents increasingly interact with external systems through tool calls--actions with real-world consequences that text outputs alone do not carry. Safety evaluations, however, overwhelmingly measure…

人工智能 · 计算机科学 2026-02-20 Arnold Cartagena , Ariane Teixeira

Safety risks of AI models have been widely studied at deployment time, such as jailbreak attacks that elicit harmful outputs. In contrast, safety risks emerging during training remain largely unexplored. Beyond explicit reward hacking that…

User models in information retrieval rest on a foundational assumption that observed behavior reveals intent. This assumption collapses when the user is an AI agent privately configured by a human operator. For any action an agent takes, a…

Healthcare conversational AI agents shouldn't be optimized only for clean benchmark accuracy in production-first regime; they must be optimized for the lived reality of patient conversations, where audio is imperfect, intent is indirect,…

We study the tendency of AI systems to deceive by constructing a realistic simulation setting of a company AI assistant. The simulated company employees provide tasks for the assistant to complete, these tasks spanning writing assistance,…

计算与语言 · 计算机科学 2024-05-06 Olli Järviniemi , Evan Hubinger

Language model alignment has become an important component of AI safety, allowing safe interactions between humans and language models, by enhancing desired behaviors and inhibiting undesired ones. It is often done by tuning the model or…

计算与语言 · 计算机科学 2025-05-28 Yotam Wolf , Noam Wies , Dorin Shteyman , Binyamin Rothberg , Yoav Levine , Amnon Shashua

As AI systems advance in capabilities, measuring their safety and alignment to human values is becoming paramount. A fast-growing field of AI research is devoted to developing such assessments. However, most current advances therein may be…

计算机与社会 · 计算机科学 2026-03-17 Max Hellrigel-Holderbaum , Edward James Young

Large language models (LLMs) are increasingly deployed to simulate human collective behaviors, yet the methodological rigor of these "AI societies" remains under-explored. Through a systematic audit of 39 recent studies, we identify six…

计算与语言 · 计算机科学 2026-04-07 Jiaxu Zhou , Jen-tse Huang , Xuhui Zhou , Man Ho Lam , Xintao Wang , Hao Zhu , Wenxuan Wang , Maarten Sap

Training large language models to follow instructions makes them perform better on a wide range of tasks and generally become more helpful. However, a perfectly helpful model will follow even the most malicious instructions and readily…

计算与语言 · 计算机科学 2024-03-20 Federico Bianchi , Mirac Suzgun , Giuseppe Attanasio , Paul Röttger , Dan Jurafsky , Tatsunori Hashimoto , James Zou

This paper presents an argument that certain AI safety measures, rather than mitigating existential risk, may instead exacerbate it. Under certain key assumptions - the inevitability of AI failure, the expected correlation between an AI…

人工智能 · 计算机科学 2024-06-04 Herman Cappelen , Josh Dever , John Hawthorne

As AI agents are increasingly adopted to collaborate on complex objectives, ensuring the security of autonomous multi-agent systems becomes crucial. We develop simulations of agents collaborating on shared objectives to study these security…

Both the general public and academic communities have raised concerns about sycophancy, the phenomenon of artificial intelligence (AI) excessively agreeing with or flattering users. Yet, beyond isolated media reports of severe consequences,…

计算机与社会 · 计算机科学 2025-10-03 Myra Cheng , Cinoo Lee , Pranav Khadpe , Sunny Yu , Dyllan Han , Dan Jurafsky

Recent and unremitting capability advances have been accompanied by calls for comprehensive, rather than patchwork, regulation of frontier artificial intelligence (AI). Approval regulation is emerging as a promising candidate. An approval…

计算机与社会 · 计算机科学 2024-08-13 Cole Salvador

The adoption of machine-learning-enabled systems in the healthcare domain is on the rise. While the use of ML in healthcare has several benefits, it also expands the threat surface of medical systems. We show that the use of ML in medical…

密码学与安全 · 计算机科学 2024-04-15 Mohammed Elnawawy , Mohammadreza Hallajiyan , Gargi Mitra , Shahrear Iqbal , Karthik Pattabiraman

Artificial intelligence (AI) holds great promise for supporting clinical trials, from patient recruitment and endpoint assessment to treatment response prediction. However, deploying AI without safeguards poses significant risks,…

机器学习 · 计算机科学 2025-10-09 Yao Chen , David Ohlssen , Aimee Readie , Gregory Ligozio , Ruvie Martin , Thibaud Coroller

We demonstrate a situation in which Large Language Models, trained to be helpful, harmless, and honest, can display misaligned behavior and strategically deceive their users about this behavior without being instructed to do so. Concretely,…

计算与语言 · 计算机科学 2024-07-16 Jérémy Scheurer , Mikita Balesni , Marius Hobbhahn