中文
相关论文

相关论文: Sustaining AI safety: Control-theoretic external i…

200 篇论文

Since the release of ChatGPT, there has been a lot of debate about whether AI systems pose an existential risk to humanity. This paper develops a general framework for thinking about the existential risk of AI systems. We analyze a two…

人工智能 · 计算机科学 2026-01-16 Herman Cappelen , Simon Goldstein , John Hawthorne

Sequential decision making using Markov Decision Process underpins many realworld applications. Both model-based and model free methods have achieved strong results in these settings. However, real-world tasks must balance reward…

机器学习 · 计算机科学 2026-04-01 Janaka Chathuranga Brahmanage , Akshat Kumar

The AI alignment problem, which focusses on ensuring that artificial intelligence (AI), including AGI and ASI, systems act according to human values, presents profound challenges. With the progression from narrow AI to Artificial General…

人工智能 · 计算机科学 2025-07-25 Alberto Hernández-Espinosa , Felipe S. Abrahão , Olaf Witkowski , Hector Zenil

To achieve reliable, robust, and safe AI systems, it is vital to implement fallback strategies when AI predictions cannot be trusted. Certifiers for neural networks are a reliable way to check the robustness of these predictions. They…

机器学习 · 计算机科学 2023-10-04 Tobias Lorenz , Marta Kwiatkowska , Mario Fritz

The rapid advancement of artificial intelligence (AI) technologies presents profound challenges to societal safety. As AI systems become more capable, accessible, and integrated into critical services, the dual nature of their potential is…

人工智能 · 计算机科学 2024-12-06 Giulio Corsi , Kyle Kilian , Richard Mallah

This paper contains the first formal proof that safety is non-compositional in the presence of conjunctive capability dependencies: two agents each individually inca- pable of reaching any forbidden capability can, when combined,…

人工智能 · 计算机科学 2026-03-20 Cosimo Spera

Effective governance of artificial intelligence (AI) requires public engagement, yet communication strategies centered on existential risk have not produced sustained mobilization. In this paper, we examine the psychological and opinion…

计算机与社会 · 计算机科学 2025-11-13 Philip Trippenbach , Isabella Scala , Jai Bhambra , Rowan Emslie

This chapter explores the symbiotic relationship between Artificial Intelligence (AI) and trust in networked systems, focusing on how these two elements reinforce each other in strategic cybersecurity contexts. AI's capabilities in data…

人工智能 · 计算机科学 2024-11-21 Yunfei Ge , Quanyan Zhu

Invention of artificial general intelligence is predicted to cause a shift in the trajectory of human civilization. In order to reap the benefits and avoid pitfalls of such powerful technology it is important to be able to control it.…

计算机与社会 · 计算机科学 2020-08-11 Roman V. Yampolskiy

With the introduction of Artificial Intelligence (AI) and related technologies in our daily lives, fear and anxiety about their misuse as well as the hidden biases in their creation have led to a demand for regulation to address such…

人工智能 · 计算机科学 2021-04-09 The Anh Han , Tom Lenaerts , Francisco C. Santos , Luis Moniz Pereira

As artificial intelligence (AI) systems become increasingly adopted across sectors, the need for robust, proactive security strategies is paramount. Traditional defensive measures often fall short against the unique and evolving threats…

密码学与安全 · 计算机科学 2025-05-13 Josh Harguess , Chris M. Ward

Robotics, automation, and related Artificial Intelligence (AI) systems have become pervasive bringing in concerns related to security, safety, accuracy, and trust. With growing dependency on physical robots that work in close proximity to…

密码学与安全 · 计算机科学 2023-02-17 Sudip Mittal , Jingdao Chen

As frontier artificial intelligence (AI) systems become more capable, it becomes more important that developers can explain why their systems are sufficiently safe. One way to do so is via safety cases: reports that make a structured…

计算机与社会 · 计算机科学 2024-10-30 Marie Davidsen Buhl , Gaurav Sett , Leonie Koessler , Jonas Schuett , Markus Anderljung

In the rapidly evolving field of cybersecurity, ensuring the reproducibility of AI-driven research is critical to maintaining the reliability and integrity of security systems. This paper addresses the reproducibility crisis within the…

机器学习 · 计算机科学 2024-12-17 Richard H. Moulton , Gary A. McCully , John D. Hastings

Security risks from AI have motivated calls for international agreements that guardrail the technology. However, even if states could agree on what rules to set on AI, the problem of verifying compliance might make these agreements…

计算机与社会 · 计算机科学 2023-04-11 Mauricio Baker

Trust is a central component of the interaction between people and AI, in that 'incorrect' levels of trust may cause misuse, abuse or disuse of the technology. But what, precisely, is the nature of trust in AI? What are the prerequisites…

人工智能 · 计算机科学 2021-01-21 Alon Jacovi , Ana Marasović , Tim Miller , Yoav Goldberg

The accelerating development and deployment of AI technologies depend on the continued ability to scale their infrastructure. This has implied increasing amounts of monetary investment and natural resources. Frontier AI applications have…

计算机与社会 · 计算机科学 2025-02-04 Eshta Bhardwaj , Rohan Alexander , Christoph Becker

Generative AI systems produce a range of risks. To ensure the safety of generative AI systems, these risks must be evaluated. In this paper, we make two main contributions toward establishing such evaluations. First, we propose a…

As AI assistants become commonplace in daily life, the demand for solutions that reduce the cost of inference without sacrificing utility is increasing. Existing work on AI sustainability frequently emphasizes hardware and software…

软件工程 · 计算机科学 2026-04-01 Madeline Jennings , Novarun Deb , Ronnie de Souza Santos

Artificial intelligence (AI) systems have become increasingly popular in many areas. Nevertheless, AI technologies are still in their developing stages, and many issues need to be addressed. Among those, the reliability of AI systems needs…

软件工程 · 计算机科学 2021-11-11 Yili Hong , Jiayi Lian , Li Xu , Jie Min , Yueyao Wang , Laura J. Freeman , Xinwei Deng