English
Related papers

Related papers: AISafetyLab: A Comprehensive Framework for AI Safe…

200 papers

While artificial intelligence (AI) is advancing rapidly and mastering increasingly complex problems with astonishing performance, the safety assurance of such systems is a major concern. Particularly in the context of safety-critical,…

Artificial Intelligence · Computer Science 2025-07-01 Lars Ullrich , Walter Zimmer , Ross Greer , Knut Graichen , Alois C. Knoll , Mohan Trivedi

Recent work has raised concerns on the risk of unintended bias in AI systems being used nowadays that can affect individuals unfairly based on race, gender or religion, among other possible characteristics. While a lot of bias metrics and…

Machine Learning · Computer Science 2019-04-30 Pedro Saleiro , Benedict Kuester , Loren Hinkson , Jesse London , Abby Stevens , Ari Anisfeld , Kit T. Rodolfa , Rayid Ghani

WalledEval is a comprehensive AI safety testing toolkit designed to evaluate large language models (LLMs). It accommodates a diverse range of models, including both open-weight and API-based ones, and features over 35 safety benchmarks…

Rising concern for the societal implications of artificial intelligence systems has inspired a wave of academic and journalistic literature in which deployed systems are audited for harm by investigators from outside the organizations…

As frontier AI models become more capable, evaluating their potential to enable cyberattacks is crucial for ensuring the safe development of Artificial General Intelligence (AGI). Current cyber evaluation efforts are often ad-hoc, lacking…

Cryptography and Security · Computer Science 2025-04-23 Mikel Rodriguez , Raluca Ada Popa , Four Flynn , Lihao Liang , Allan Dafoe , Anna Wang

Last decade has seen major improvements in the performance of artificial intelligence which has driven wide-spread applications. Unforeseen effects of such mass-adoption has put the notion of AI safety into the public eye. AI safety is a…

Computers and Society · Computer Science 2020-07-10 Mislav Juric , Agneza Sandic , Mario Brcic

Recent advances in machine learning, particularly the emergence of foundation models, are leading to new opportunities to develop technology-based solutions to societal problems. However, the reasoning and inner workings of today's complex…

Computers and Society · Computer Science 2025-07-01 Rajeev Alur , Greg Durrett , Hadas Kress-Gazit , Corina Păsăreanu , René Vidal

Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general intelligence, or systems used in safety-critical contexts. In…

The popularity of large language models (LLMs) continues to grow, and LLM-based assistants have become ubiquitous. Information security awareness (ISA) is an important yet underexplored area of LLM safety. ISA encompasses LLMs' security…

Cryptography and Security · Computer Science 2026-03-23 Ofir Cohen , Gil Ari Agmon , Asaf Shabtai , Rami Puzis

Artificial Intelligence (AI) has emerged as a key technology, driving advancements across a range of applications. Its integration into modern autonomous systems requires assuring safety. However, the challenge of assuring safety in systems…

Machine Learning · Computer Science 2024-12-19 Ronald Schnitzer , Lennart Kilian , Simon Roessner , Konstantinos Theodorou , Sonja Zillner

The exposure of security vulnerabilities in safety-aligned language models, e.g., susceptibility to adversarial attacks, has shed light on the intricate interplay between AI safety and AI security. Although the two disciplines now come…

As AI systems proliferate in society, the AI community is increasingly preoccupied with the concept of AI Safety, namely the prevention of failures due to accidents that arise from an unanticipated departure of a system's behavior from…

Computers and Society · Computer Science 2024-01-23 Inioluwa Deborah Raji , Roel Dobbe

The United States Department of Defense (DOD) looks to accelerate the development and deployment of AI capabilities across a wide spectrum of defense applications to maintain strategic advantages. However, many common features of AI…

Artificial Intelligence · Computer Science 2025-04-25 Ariel S. Kapusta , David Jin , Peter M. Teague , Robert A. Houston , Jonathan B. Elliott , Grace Y. Park , Shelby S. Holdren

AI and its relevant technologies, including machine learning, deep learning, chatbots, virtual assistants, and others, are currently undergoing a profound transformation of development and organizational processes within companies.…

Cryptography and Security · Computer Science 2024-12-11 Tingting Bi , Guangsheng Yu , Qin Wang

Medical Large Language Models (LLMs) are increasingly deployed for clinical decision support across diverse specialties, yet systematic evaluation of their robustness to adversarial misuse and privacy leakage remains inaccessible to most…

Cryptography and Security · Computer Science 2025-12-10 Jinghao Wang , Ping Zhang , Carter Yagemann

Artificial Intelligence (AI) systems are now an integral part of multiple industries. In clinical research, AI supports automated adverse event detection in clinical trials, patient eligibility screening for protocol enrollment, and data…

Artificial Intelligence · Computer Science 2025-12-10 Laxmiraju Kandikatla , Branislav Radeljic

Frontier artificial intelligence (AI) systems pose increasing risks to society, making it essential for developers to provide assurances about their safety. One approach to offering such assurances is through a safety case: a structured,…

Computers and Society · Computer Science 2024-11-14 Arthur Goemans , Marie Davidsen Buhl , Jonas Schuett , Tomek Korbak , Jessica Wang , Benjamin Hilton , Geoffrey Irving

We present CAISAR, an open-source platform under active development for the characterization of AI systems' robustness and safety. CAISAR provides a unified entry point for defining verification problems by using WhyML, the mature and…

Artificial Intelligence · Computer Science 2022-06-22 Julien Girard-Satabin , Michele Alberti , François Bobot , Zakaria Chihani , Augustin Lemesle

The rapid growth of Artificial Intelligence (AI) has underscored the urgent need for responsible AI practices. Despite increasing interest, a comprehensive AI risk assessment toolkit remains lacking. This study introduces our Responsible AI…

Computers and Society · Computer Science 2025-01-23 Sung Une Lee , Harsha Perera , Yue Liu , Boming Xia , Qinghua Lu , Liming Zhu , Olivier Salvado , Jon Whittle

Embodied AI systems, including robots and autonomous vehicles, are increasingly integrated into real-world applications, where they encounter a range of vulnerabilities stemming from both environmental and system-level factors. These…

Cryptography and Security · Computer Science 2025-02-26 Wenpeng Xing , Minghao Li , Mohan Li , Meng Han