中文
相关论文

相关论文: How Well Do Models Follow Their Constitutions?

200 篇论文

Prominent AI companies are producing 'safety frameworks' as a type of voluntary self-governance. These statements purport to establish risk thresholds and safety procedures for the development and deployment of highly capable AI.…

计算机与社会 · 计算机科学 2025-10-14 Sam Coggins , Alexander K. Saeri , Katherine A. Daniell , Lorenn P. Ruster , Jessie Liu , Jenny L. Davis

AI leaders and safety reports increasingly warn that advances in model reasoning may enable biological misuse, including by low-expertise users, while major labs describe safeguards as expanding but still evolving rather than settled. This…

计算机与社会 · 计算机科学 2026-04-24 Michael Richter

Foundation models, i.e. large neural networks pre-trained on large text corpora, have revolutionized NLP. They can be instructed directly (e.g. (arXiv:2005.14165)) - this is called hard prompting - and they can be tuned using very little…

计算与语言 · 计算机科学 2023-06-13 Sid Mittal , Vineet Gupta , Frederick Liu , Mukund Sundararajan

In many legal processes being able to action on the concrete implication of a legal question can be valuable to automating human review or signalling certain conditions (e.g., alerts around automatic renewal). To support such tasks, we…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Adam Roegiest , Radha Chitta , Jonathan Donnelly , Maya Lash , Alexandra Vtyurina , François Longtin

In safety-critical applications such as medical imaging and autonomous driving, where decisions have profound implications for patient health and road safety, it is imperative to maintain both high adversarial robustness to protect against…

机器学习 · 计算机科学 2024-05-16 Ziquan Liu , Yufei Cui , Yan Yan , Yi Xu , Xiangyang Ji , Xue Liu , Antoni B. Chan

This paper leverages insights from Alignment Theory (AT) research, which primarily focuses on the potential pitfalls of technical alignment in Artificial Intelligence, to critically examine the European Union's Artificial Intelligence Act…

计算机与社会 · 计算机科学 2024-10-29 Alejandro Tlaie

This article describes how technical infrastructure developed by the nonprofit OpenMined enables external scrutiny of AI systems without compromising sensitive information. Independent external scrutiny of AI systems provides crucial…

计算机与社会 · 计算机科学 2025-02-11 Kendrea Beers , Helen Toner

Several jurisdictions are starting to regulate frontier artificial intelligence (AI) systems, i.e. general-purpose AI systems that match or exceed the capabilities present in the most advanced systems. To reduce risks from these systems,…

计算机与社会 · 计算机科学 2025-08-27 Jonas Schuett , Markus Anderljung , Alexis Carlier , Leonie Koessler , Ben Garfinkel

Frontier AI models -- highly capable foundation models at the cutting edge of AI development -- may pose severe risks to public safety, human rights, economic stability, and societal value in the coming years. These risks could arise from…

计算机与社会 · 计算机科学 2025-03-11 Deepika Raman , Nada Madkour , Evan R. Murphy , Krystal Jackson , Jessica Newman

Industrial teams often deploy large language model features before stable regression or model selection evaluation exists. We present a reusable evaluation system for AI meeting summaries that combines structured ground-truth (GT)…

人工智能 · 计算机科学 2026-05-14 Philip Zhong , Don Wang , Jason Zhang

Rising concern for the societal implications of artificial intelligence systems has inspired a wave of academic and journalistic literature in which deployed systems are audited for harm by investigators from outside the organizations…

The last decade has witnessed the proliferation of Deep Learning models in many applications, achieving unrivaled levels of predictive performance. Unfortunately, the black-box nature of Deep Learning models has posed unanswered questions…

机器学习 · 计算机科学 2020-03-26 Alejandro Barredo-Arrieta , Javier Del Ser

Recent work has developed optimization procedures to find token sequences, called adversarial triggers, which can elicit unsafe responses from aligned language models. These triggers are believed to be highly transferable, i.e., a trigger…

计算与语言 · 计算机科学 2025-04-10 Nicholas Meade , Arkil Patel , Siva Reddy

AI systems that output their reasoning in natural language offer an opportunity for safety -- we can \emph{monitor} their chain of thought (CoT) for undesirable reasoning, such as the pursuit of harmful objectives. However, the extent to…

人工智能 · 计算机科学 2025-12-10 Matt MacDermott , Qiyao Wei , Rada Djoneva , Francis Rhys Ward

Technical standards, or simply standards, are established documented guidelines and rules that facilitate the interoperability, quality, and accuracy of systems and processes. In recent years, we have witnessed an emerging paradigm shift…

计算机与社会 · 计算机科学 2025-03-10 Joseph Marvin Imperial , Matthew D. Jones , Harish Tayyar Madabushi

Auditing mechanisms for differential privacy use probabilistic means to empirically estimate the privacy level of an algorithm. For private machine learning, existing auditing mechanisms are tight: the empirical privacy estimate (nearly)…

As generative AI (GenAI) is increasingly applied in persona development to represent real users, understanding the implications and limitations of this technology is essential for establishing robust practices. This scoping review analyzes…

人机交互 · 计算机科学 2026-04-20 Danial Amin , Joni Salminen , Farhan Ahmed , Sonja M. H. Tervola , Sankalp Sethi , Bernard J. Jansen

Trustworthy capability evaluations are crucial for ensuring the safety of AI systems, and are becoming a key component of AI regulation. However, the developers of an AI system, or the AI system itself, may have incentives for evaluations…

人工智能 · 计算机科学 2025-02-10 Teun van der Weij , Felix Hofstätter , Ollie Jaffe , Samuel F. Brown , Francis Rhys Ward

The legal field already uses various large language models (LLMs) in actual applications, but their quantitative performance and reasons for it are underexplored. We evaluated several open-source and proprietary LLMs -- including…

计算机与社会 · 计算机科学 2025-09-12 Bhakti Khera , Rezvan Alamian , Pascal A. Scherz , Stephan M. Goetz

Generative Artificial Intelligence (AI) models such as OpenAI's ChatGPT have the potential to revolutionize Statistical Process Control (SPC) practice, learning, and research. However, these tools are in the early stages of development and…

机器学习 · 计算机科学 2023-06-19 Fadel M. Megahed , Ying-Ju Chen , Joshua A. Ferris , Sven Knoth , L. Allison Jones-Farmer