English
Related papers

Related papers: Evaluating the Critical Risks of Amazon's Nova Pre…

200 papers

Multimodal foundation models (MMFMs) play a crucial role in various applications, including autonomous driving, healthcare, and virtual assistants. However, several studies have revealed vulnerabilities in these models, such as generating…

Safe learning is central to AI-enabled robots where a single failure may lead to catastrophic results. Barrier-based method is one of the dominant approaches for safe robot learning. However, this method is not scalable, hard to train, and…

Machine Learning · Computer Science 2024-06-21 Wei Xiao , Tsun-Hsuan Wang , Daniela Rus

Chai empowers users to create and interact with customized chatbots, offering unique and engaging experiences. Despite the exciting prospects, the work recognizes the inherent challenges of a commitment to modern safety standards.…

Artificial Intelligence · Computer Science 2023-06-06 Xiaoding Lu , Aleksey Korshuk , Zongyi Liu , William Beauchamp

This paper contributes to the nascent debate around safety cases for frontier AI systems. Safety cases are structured, defensible arguments that a system is acceptably safe to deploy in a given context. Historically, they have been used in…

Computers and Society · Computer Science 2026-03-11 Shaun Feakins , Ibrahim Habli , Phillip Morgan

As foundation models grow in both popularity and capability, researchers have uncovered a variety of ways that the models can pose a risk to the model's owner, user, or others. Despite the efforts of measuring these risks via benchmarks and…

Cryptography and Security · Computer Science 2025-06-04 David Piorkowski , Michael Hind , John Richards , Jacquelyn Martino

This year, jurisdictions worldwide, including the United States, the European Union, the United Kingdom, and China, are set to enact or revise laws governing frontier AI. Their efforts largely rely on the assumption that increasing model…

Computers and Society · Computer Science 2025-02-25 Nicholas A. Caputo

All of the frontier AI companies have published safety frameworks where they define capability thresholds and risk mitigations that determine how they will safely develop and deploy their models. Adoption of systematic approaches to risk…

Computers and Society · Computer Science 2025-06-03 Simon Mylius

Security evaluations inherently depend on stable identifiers. Any finding, audit, or regulatory decision must remain attached to the specific artifact it pertains to. Continuously updated artificial intelligence systems violate this core…

Cryptography and Security · Computer Science 2026-05-26 Dan Ristea , Vasilios Mavroudis

Open-domain video platforms offer rich, personalized content that could support health, caregiving, and educational applications, but their engagement-optimized recommendation algorithms can expose vulnerable users to inappropriate or…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Wenzheng Zhao , Madhava Kalyan Gadiputi , Fengpei Yuan

Frontier AI systems are rapidly advancing in their capabilities to persuade, deceive, and influence human behaviour, with current models already demonstrating human-level persuasion and strategic deception in specific contexts. Humans are…

Artificial Intelligence · Computer Science 2025-07-18 Rishane Dassanayake , Mario Demetroudi , James Walpole , Lindley Lentati , Jason R. Brown , Edward James Young

Safety and responsibility evaluations of advanced AI models are a critical but developing field of research and practice. In the development of Google DeepMind's advanced AI models, we innovated on and applied a broad set of approaches to…

Safety-critical environments are inherently dynamic. Distribution shifts, emerging vulnerabilities, and evolving requirements demand continuous updates to machine learning models. Yet even benign parameter updates can have unintended…

Machine Learning · Computer Science 2026-03-19 Leo Elmecker-Plakolm , Pierre Fasterling , Philip Sosnin , Calvin Tsay , Matthew Wicker

Continued adoption of agricultural robots postulates the farmer's trust in the reliability, robustness and safety of the new technology. This motivates our work on safety assurance of agricultural robots, particularly their ability to…

Robotics · Computer Science 2025-06-25 Mustafa Adam , Kangfeng Ye , David A. Anisi , Ana Cavalcanti , Jim Woodcock , Robert Morris

Vulnerability of Frontier language models to misuse and jailbreaks has prompted the development of safety measures like filters and alignment training in an effort to ensure safety through robustness to adversarially crafted prompts. We…

Cryptography and Security · Computer Science 2024-10-31 David Glukhov , Ziwen Han , Ilia Shumailov , Vardan Papyan , Nicolas Papernot

Ensuring both performance and safety is critical for autonomous systems operating in real-world environments. While safety filters such as Control Barrier Functions (CBFs) enforce constraints by modifying nominal controllers in real time,…

Systems and Control · Electrical Eng. & Systems 2025-07-21 Aditya Singh , Aastha Mishra , Manan Tayal , Shishir Kolathaya , Pushpak Jagtap

Frontier AI systems are being adopted across Africa, yet most AI safety evaluations are designed and validated in Western environments. In this paper, we argue that the portability gap can leave Africa-centric pathways to severe harm…

As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overlooked in traditional AI risk assessment frameworks. This…

Cryptography and Security · Computer Science 2024-10-21 Aviral Srivastava , Sourav Panda

Frontier AI both amplifies existing risks and introduces qualitatively novel challenges. Not only is there a notable lack of stable scientific consensus resulting from the rapid pace of technological change, but emerging frontier AI safety…

Evaluations of dangerous AI capabilities are important for managing catastrophic risks. Public transparency into these evaluations - including what they test, how they are conducted, and how their results inform decisions - is crucial for…

Computers and Society · Computer Science 2025-09-04 Tegan McCaslin , Jide Alaga , Samira Nedungadi , Seth Donoughe , Tom Reed , Rishi Bommasani , Chris Painter , Luca Righetti

AI model documentation is fragmented across platforms and inconsistent in structure, preventing policymakers, auditors, and users from reliably assessing safety claims, data provenance, and version-level changes. We analyzed documentation…

Artificial Intelligence · Computer Science 2025-12-16 Akhmadillo Mamirov , Faiaz Azmain , Hanyu Wang