English
Related papers

Related papers: An Approach to Technical AGI Safety and Security

200 papers

The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural language understanding and generation. However, the increasing…

Artificial Intelligence · Computer Science 2024-12-25 Dan Shi , Tianhao Shen , Yufei Huang , Zhigen Li , Yongqi Leng , Renren Jin , Chuang Liu , Xinwei Wu , Zishan Guo , Linhao Yu , Ling Shi , Bojian Jiang , Deyi Xiong

As AI systems become more advanced, concerns about large-scale risks from misuse or accidents have grown. This report analyzes the technical research into safe AI development being conducted by three leading AI companies: Anthropic, Google…

Computers and Society · Computer Science 2024-09-26 Oscar Delaney , Oliver Guest , Zoe Williams

Recent advances in artificial general intelligence (AGI), particularly large language models and creative image generation systems have demonstrated impressive capabilities on diverse tasks spanning the arts and humanities. However, the…

Artificial Intelligence · Computer Science 2023-10-31 Zhengliang Liu , Yiwei Li , Qian Cao , Junwen Chen , Tianze Yang , Zihao Wu , John Hale , John Gibbs , Khaled Rasheed , Ninghao Liu , Gengchen Mai , Tianming Liu

AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm. Untrusted monitoring -- using one untrusted model to oversee another -- is one…

Artificial intelligence (AI) promises immense benefits across sectors, yet also poses risks from dual-use potentials, biases, and unintended behaviors. This paper reviews emerging issues with opaque and uncontrollable AI systems and…

Artificial Intelligence · Computer Science 2023-08-29 Alexander J. Titus , Adam H. Russell

We outline the principles of classical assurance for computer-based systems that pose significant risks. We then consider application of these principles to systems that employ Artificial Intelligence (AI) and Machine Learning (ML). A key…

Artificial Intelligence · Computer Science 2025-06-04 Robin Bloomfield , John Rushby

In this 4-page manuscript we discuss the problem of long-term AI Safety from a Software Engineering (SE) research viewpoint. We briefly summarize long-term AI Safety, and the challenge of avoiding harms from AI as systems meet or exceed…

Software Engineering · Computer Science 2023-09-01 David Gros , Prem Devanbu , Zhou Yu

Humanity is progressing towards automated product development, a trend that promises faster creation of better products and thus the acceleration of technological progress. However, increasing reliance on non-human agents for this process…

Computers and Society · Computer Science 2025-06-03 Jan Göpfert , Jann M. Weinand , Patrick Kuckertz , Noah Pflugradt , Jochen Linßen

AI-based systems have been used widely across various industries for different decisions ranging from operational decisions to tactical and strategic ones in low- and high-stakes contexts. Gradually the weaknesses and issues of these…

Human-Computer Interaction · Computer Science 2022-01-13 Morteza Saberi

This paper summarizes the most cogent advantages and risks associated with Artificial Intelligence from an in-depth review of the literature. Then the authors synthesize the salient risk-related models currently being used in AI, technology…

Computers and Society · Computer Science 2024-06-19 Richard Fulton , Diane Fulton , Nate Hayes , Susan Kaplan

Advanced AI models hold the promise of tremendous benefits for humanity, but society needs to proactively manage the accompanying risks. In this paper, we focus on what we term "frontier AI" models: highly capable foundation models that…

While Artificial Intelligence (AI) is not a new field, recent developments, especially with the release of generative tools like ChatGPT, have brought it to the forefront of the minds of industry workers and academic folk alike. There is…

Computers and Society · Computer Science 2025-07-22 Prerana Khatiwada , Grace Donaher , Jasymyn Navarro , Lokesh Bhatta

Artificial intelligence (AI) has been advancing at a fast pace and it is now poised for deployment in a wide range of applications, such as autonomous systems, medical diagnosis and natural language processing. Early adoption of AI…

Machine Learning · Computer Science 2023-09-21 Marta Kwiatkowska , Xiyue Zhang

Generative AI (GenAI) is a powerful technology poised to reshape Trust & Safety. While misuse by attackers is a growing concern, its defensive capacity remains underexplored. This paper examines these effects through a qualitative study…

Human-Computer Interaction · Computer Science 2026-04-24 Patrick Gage Kelley , Steven Rousso-Schindler , Renee Shelby , Kurt Thomas , Allison Woodruff

This article discusses some trends and concepts in developing new generation of future Artificial General Intelligence (AGI) systems which relate to complex facets and different types of human intelligence, especially social, emotional,…

Artificial Intelligence · Computer Science 2020-12-14 Andrzej Cichocki , Alexander P. Kuleshov

In an era where the Internet of Things (IoT) intersects increasingly with generative Artificial Intelligence (AI), this article scrutinizes the emergent security risks inherent in this integration. We explore how generative AI drives…

Cryptography and Security · Computer Science 2024-04-02 Honghui Xu , Yingshu Li , Olusesi Balogun , Shaoen Wu , Yue Wang , Zhipeng Cai

AI agents have been boosted by large language models. AI agents can function as intelligent assistants and complete tasks on behalf of their users with access to tools and the ability to execute commands in their environments. Through…

Cryptography and Security · Computer Science 2024-12-19 Yifeng He , Ethan Wang , Yuyang Rong , Zifei Cheng , Hao Chen

With the advent of the digital era, every day-to-day task is automated due to technological advances. However, technology has yet to provide people with enough tools and safeguards. As the internet connects more-and-more devices around the…

Cryptography and Security · Computer Science 2022-09-28 Abhilash Chakraborty , Anupam Biswas , Ajoy Kumar Khan

Alignment of artificial intelligence (AI) encompasses the normative problem of specifying how AI systems should act and the technical problem of ensuring AI systems comply with those specifications. To date, AI alignment has generally…

Artificial intelligence (AI) is reshaping society, from video generation to medical diagnosis, coding agents to autonomous vehicles. Yet researchers, policymakers, and technology companies lack shared terminology for discussing AI risks.…