English
Related papers

Related papers: Robust AI Security and Alignment: A Sisyphean Ende…

200 papers

Recent progress in artificial intelligence (AI) using deep learning techniques has triggered its wide-scale use across a broad range of applications. These systems can already perform tasks such as natural language processing of voice and…

Computers and Society · Computer Science 2019-10-29 P. Santhanam , Eitan Farchi , Victor Pankratius

Artificial Intelligence (AI) has the opportunity to revolutionize the way the United States Department of Defense (DoD) and Intelligence Community (IC) address the challenges of evolving threats, data deluge, and rapid courses of action.…

Artificial Intelligence · Computer Science 2019-05-10 Vijay Gadepally , Justin Goodwin , Jeremy Kepner , Albert Reuther , Hayley Reynolds , Siddharth Samsi , Jonathan Su , David Martinez

Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose existential risks. This paper reviews the…

Computers and Society · Computer Science 2023-10-30 Rose Hadshar

As AI systems increasingly influence critical decisions, they face threats that exploit reasoning mechanisms rather than technical infrastructure. We present a framework for cognitive cybersecurity, a systematic protection of AI reasoning…

Cryptography and Security · Computer Science 2025-08-25 Yuksel Aydin

In recent years there has been a tremendous surge in the general capabilities of AI systems, mainly fuelled by training foundation models on internetscale data. Nevertheless, the creation of openended, ever self-improving AI remains…

5G networks offer exceptional reliability and availability, ensuring consistent performance and user satisfaction. Yet they might still fail when confronted with the unexpected. A resilient system is able to adapt to real-world complexity,…

Information Theory · Computer Science 2025-10-27 Bho Matthiesen , Armin Dekorsy , Petar Popovski

As AI models scale to billions of parameters and operate with increasing autonomy, ensuring their safe, reliable operation demands engineering-grade security and assurance frameworks. This paper presents an enterprise-level, risk-aware,…

Cryptography and Security · Computer Science 2025-05-13 Krti Tallam

This thesis investigates three areas targeted at improving the reliability of machine learning; fairness in machine learning, strategic classification, and algorithmic robustness. Each of these domains has special properties or structure…

Machine Learning · Computer Science 2024-08-30 Kevin Stangl

As we seek to deploy machine learning models beyond virtual and controlled domains, it is critical to analyze not only the accuracy or the fact that it works most of the time, but if such a model is truly robust and reliable. This paper…

Machine Learning · Computer Science 2020-07-07 Samuel Henrique Silva , Peyman Najafirad

The governance of open-weight artificial intelligence (AI) models has been framed as a binary choice: openness as risk, restriction as safety. This paper challenges that framing, arguing that access restrictions, without governed…

Computers and Society · Computer Science 2026-04-21 Vinicius Santana Gomes

Success in the quest for artificial intelligence has the potential to bring unprecedented benefits to humanity, and it is therefore worthwhile to investigate how to maximize these benefits while avoiding potential pitfalls. This article…

Artificial Intelligence · Computer Science 2016-02-15 Stuart Russell , Daniel Dewey , Max Tegmark

As AI becomes more capable, we entrust it with more general and consequential tasks. The risks from failure grow more severe with increasing task scope. It is therefore important to understand how extremely capable AI models will fail: Will…

Artificial Intelligence · Computer Science 2026-04-13 Alexander Hägele , Aryo Pradipta Gema , Henry Sleight , Ethan Perez , Jascha Sohl-Dickstein

Incompleteness theorems of Godel, Turing, Chaitin, and Algorithmic Information Theory have profound epistemological implications. Incompleteness limits our ability to ever understand every observable phenomenon in the universe.…

General Literature · Computer Science 2016-02-26 Gary R. Prok

Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant…

The increasing capabilities of artificial intelligence (AI) systems make it ever more important that we interpret their internals to ensure that their intentions are aligned with human values. Yet there is reason to believe that misaligned…

Machine Learning · Computer Science 2022-12-23 Lee Sharkey

We formalize AI alignment as a multi-objective optimization problem called $\langle M,N,\varepsilon,\delta\rangle$-agreement, in which a set of $N$ agents (including humans) must reach approximate ($\varepsilon$) agreement across $M$…

Artificial Intelligence · Computer Science 2025-11-20 Aran Nayebi

Large language models (LLMs) have become increasingly sophisticated, leading to widespread deployment in sensitive applications where safety and reliability are paramount. However, LLMs have inherent risks accompanying them, including bias,…

Cryptography and Security · Computer Science 2024-06-21 Suriya Ganesh Ayyamperumal , Limin Ge

The rapid development of AI systems poses unprecedented risks, including loss of control, misuse, geopolitical instability, and concentration of power. To navigate these risks and avoid worst-case outcomes, governments may proactively…

Artificial Intelligence · Computer Science 2025-07-15 Peter Barnett , Aaron Scher , David Abecassis

As LLMs grow more powerful, their most profound achievement may be recognising when to say "I don't know". Existing studies on LLM self-knowledge have been largely constrained by human-defined notions of feasibility, often neglecting the…

Computation and Language · Computer Science 2025-09-16 Sahil Kale , Vijaykant Nadadur

We present our Balanced, Integrated and Grounded (BIG) argument for assuring the safety of AI systems. The BIG argument adopts a whole-system approach to constructing a safety case for AI systems of varying capability, autonomy and…

Computers and Society · Computer Science 2025-04-01 Ibrahim Habli , Richard Hawkins , Colin Paterson , Philippa Ryan , Yan Jia , Mark Sujan , John McDermid