中文
相关论文

相关论文: Responsible Reporting for Frontier AI Development

200 篇论文

For AI technology to fulfill its full promises, we must have effective means to ensure Responsible AI behavior and curtail potential irresponsible use, e.g., in areas of privacy protection, human autonomy, robustness, and prevention of…

计算机与社会 · 计算机科学 2022-05-13 Wenjing Chu

The development of privacy-enhancing technologies has made immense progress in reducing trade-offs between privacy and performance in data exchange and analysis. Similar tools for structured transparency could be useful for AI governance by…

人工智能 · 计算机科学 2023-03-22 Emma Bluemke , Tantum Collins , Ben Garfinkel , Andrew Trask

Safety and responsibility evaluations of advanced AI models are a critical but developing field of research and practice. In the development of Google DeepMind's advanced AI models, we innovated on and applied a broad set of approaches to…

Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavior could be very hard to detect, and if present in future…

All types of research, development, and policy work can have unintended, adverse consequences - work in responsible artificial intelligence (RAI), ethical AI, or ethics in AI is no exception.

人工智能 · 计算机科学 2023-11-21 Alexandra Olteanu , Michael Ekstrand , Carlos Castillo , Jina Suh

Artificial intelligence (AI) and machine learning (ML) have made tremendous advancements in the past decades. From simple recommendation systems to more complex tumor identification systems, AI/ML systems have been utilized in a plethora of…

计算机与社会 · 计算机科学 2025-02-07 Atul Rawal , Katie Johnson , Curtis Mitchell , Michael Walton , Diamond Nwankwo

The rapid development of AI systems poses unprecedented risks, including loss of control, misuse, geopolitical instability, and concentration of power. To navigate these risks and avoid worst-case outcomes, governments may proactively…

人工智能 · 计算机科学 2025-07-15 Peter Barnett , Aaron Scher , David Abecassis

In the past few years, several large companies have published ethical principles of Artificial Intelligence (AI). National governments, the European Commission, and inter-governmental organizations have come up with requirements to ensure…

计算机与社会 · 计算机科学 2020-05-06 Richard Benjamins

The rapid development of artificial intelligence (AI) has led to increasing concerns about the capability of AI systems to make decisions and behave responsibly. Responsible AI (RAI) refers to the development and use of AI systems that…

软件工程 · 计算机科学 2023-05-25 Boming Xia , Qinghua Lu , Harsha Perera , Liming Zhu , Zhenchang Xing , Yue Liu , Jon Whittle

The EU Artificial Intelligence (AI) Act directs businesses to assess their AI systems to ensure they are developed in a way that is human-centered and trustworthy. The rapid adoption of AI in the industry has outpaced ethical evaluation…

计算机与社会 · 计算机科学 2025-09-30 Louise McCormack , Diletta Huyskes , Dave Lewis , Malika Bendechache

Artificial intelligence (AI) has been clearly established as a technology with the potential to revolutionize fields from healthcare to finance - if developed and deployed responsibly. This is the topic of responsible AI, which emphasizes…

人工智能 · 计算机科学 2023-12-05 Stephanie Baker , Wei Xiang

Responsible Artificial Intelligence (AI) proposes a framework that holds all stakeholders involved in the development of AI to be responsible for their systems. It, however, fails to accommodate the possibility of holding AI responsible per…

计算机与社会 · 计算机科学 2020-04-27 Gabriel Lima , Meeyoung Cha

Governments, industry, and other actors involved in governing AI technologies around the world agree that, while AI offers tremendous promise to benefit the world, appropriate guardrails are required to mitigate risks. Global institutions,…

计算机与社会 · 计算机科学 2024-09-18 A. Leone De Castris , C. Thomas

Risk reporting is essential for documenting AI models, yet only 14% of model cards mention risks, out of which 96% copying content from a small set of cards, leading to a lack of actionable insights. Existing proposals for improving model…

软件工程 · 计算机科学 2025-04-15 Pooja S. B. Rao , Sanja Šćepanović , Ke Zhou , Edyta Paulina Bogucka , Daniele Quercia

As Artificial Intelligence (AI) becomes integral to business operations, integrating Responsible AI (RAI) within Environmental, Social, and Governance (ESG) frameworks is essential for ethical and sustainable AI deployment. This study…

计算机与社会 · 计算机科学 2024-09-18 Harsha Perera , Sung Une Lee , Yue Liu , Boming Xia , Qinghua Lu , Liming Zhu , Jessica Cairns , Moana Nottage

Every AI system is deployed by a human organization. In high risk applications, the combined human plus AI system must function as a high-reliability organization in order to avoid catastrophic errors. This short note reviews the properties…

人工智能 · 计算机科学 2018-11-28 Thomas G. Dietterich

We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through scheming. Scheming is a potential threat model where AI…

Artificial intelligence (AI) and Machine Learning (ML) have moved from research and pilot projects into everyday business operations, with generative AI accelerating adoption across processes, products, and services. This paper introduces…

计算机与社会 · 计算机科学 2026-02-17 Stephan Sandfuchs , Diako Farooghi , Janis Mohr , Sarah Grewe , Markus Lemmen , Jörg Frochte

The downstream use cases, benefits, and risks of AI systems depend significantly on the access afforded to the system, and to whom. However, the downstream implications of different access styles are not well understood, making it difficult…

计算机与社会 · 计算机科学 2024-12-03 Edward Kembery , Ben Bucknall , Morgan Simpson

The rapid development of Artificial Intelligence (AI) technology has enabled the deployment of various systems based on it. However, many current AI systems are found vulnerable to imperceptible attacks, biased against underrepresented…

人工智能 · 计算机科学 2022-05-27 Bo Li , Peng Qi , Bo Liu , Shuai Di , Jingen Liu , Jiquan Pei , Jinfeng Yi , Bowen Zhou