English
Related papers

Related papers: Model evaluation for extreme risks

200 papers

Large language models (LLMs) are demonstrating increasing prowess in cybersecurity applications, creating creating inherent risks alongside their potential for strengthening defenses. In this position paper, we argue that current efforts to…

Cryptography and Security · Computer Science 2025-02-04 Kamilė Lukošiūtė , Adam Swanda

Sufficiently capable models could subvert human oversight and decision-making in important contexts. For example, in the context of AI development, models could covertly sabotage efforts to evaluate their own dangerous capabilities, to…

Frontier artificial intelligence (AI) systems could pose increasing risks to public safety and security. But what level of risk is acceptable? One increasingly popular approach is to define capability thresholds, which describe AI…

Computers and Society · Computer Science 2024-06-24 Leonie Koessler , Jonas Schuett , Markus Anderljung

This paper examines the critical challenges and potential solutions for conducting secure and effective external evaluations of general-purpose AI (GPAI) models. With the exponential growth in size, capability, reach and accompanying risk…

Computers and Society · Computer Science 2025-03-14 Alejandro Tlaie , Jimmy Farrell

Evaluations of large language model (LLM) risks and capabilities are increasingly being incorporated into AI risk management and governance frameworks. Currently, most risk evaluations are conducted by designing inputs that elicit harmful…

Model-based engineering promises to boost productivity and quality of complex systems development. In the context of safety-critical systems, a traditionally highly regulated and conservative domain, the use of models gained importance in…

Software Engineering · Computer Science 2021-06-07 Marc Zeller , Daniel Ratiu , Kai Hoefig

Adversarial attacks pose a severe risk to AI systems used in healthcare, capable of misleading models into dangerous misclassifications that can delay treatments or cause misdiagnoses. These attacks, often imperceptible to human perception,…

Machine Learning · Computer Science 2025-10-29 Alyssa Gerhart , Balaji Iyangar

As artificial intelligence (AI) models are scaled up, new capabilities can emerge unintentionally and unpredictably, some of which might be dangerous. In response, dangerous capabilities evaluations have emerged as a new risk assessment…

Computers and Society · Computer Science 2023-10-03 Jide Alaga , Jonas Schuett

As frontier AI models become more capable, evaluating their potential to enable cyberattacks is crucial for ensuring the safe development of Artificial General Intelligence (AGI). Current cyber evaluation efforts are often ad-hoc, lacking…

Cryptography and Security · Computer Science 2025-04-23 Mikel Rodriguez , Raluca Ada Popa , Four Flynn , Lihao Liang , Allan Dafoe , Anna Wang

This position paper argues for two claims regarding AI testing and evaluation. First, to remain informative about deployment behaviour, evaluations need account for the possibility that AI systems understand their circumstances and reason…

Computer Science and Game Theory · Computer Science 2025-08-22 Vojtech Kovarik , Eric Olav Chen , Sami Petersen , Alexis Ghersengorin , Vincent Conitzer

AI models that predict the future behavior of a system (a.k.a. predictive AI models) are central to intelligent decision-making. However, decision-making using predictive AI models often results in suboptimal performance. This is primarily…

Artificial Intelligence · Computer Science 2025-01-13 Akhil S Anand , Shambhuraj Sawant , Dirk Reinhardt , Sebastien Gros

Conventional AI evaluation approaches concentrated within the AI stack exhibit systemic limitations for exploring, navigating and resolving the human and societal factors that play out in real world deployment such as in education, finance,…

Generative AI models are capable of performing a wide variety of tasks that have traditionally required creativity and human understanding. During training, they learn patterns from existing data and can subsequently generate new content…

Artificial intelligence risks are multidimensional in nature, as the same risk scenarios may have legal, operational, and financial risk dimensions. With the emergence of new AI regulations, the state of the art of artificial intelligence…

Computers and Society · Computer Science 2025-09-24 Luis Enriquez Alvarez

Artificial Intelligence has made a significant contribution to autonomous vehicles, from object detection to path planning. However, AI models require a large amount of sensitive training data and are usually computationally intensive to…

Cryptography and Security · Computer Science 2021-09-13 Shanthi Lekkala , Tanya Motwani , Manojkumar Parmar , Amit Phadke

Machine learning techniques are effective for building predictive models because they identify patterns in large datasets. Development of a model for complex real-life problems often stop at the point of publication, proof of concept or…

Evaluating the safety of AI Systems is a pressing concern for organizations deploying them. In addition to the societal damage done by the lack of fairness of those systems, deployers are concerned about the legal repercussions and the…

Artificial Intelligence (AI) Safety Institutes and governments worldwide are deciding whether they evaluate advanced AI themselves, support a private evaluation ecosystem or do both. Evaluation regimes have been established in a wide range…

Computers and Society · Computer Science 2025-08-06 Merlin Stein , Milan Gandhi , Theresa Kriecherbauer , Amin Oueslati , Robert Trager

Embedding artificial intelligence into systems introduces significant challenges to modern engineering practices. Hazard analysis tools and processes have not yet been adequately adapted to the new paradigm. This paper describes initial…

Software Engineering · Computer Science 2022-03-30 Nikolas Martelaro , Carol J. Smith , Tamara Zilovic

Adversarial robustness studies the worst-case performance of a machine learning model to ensure safety and reliability. With the proliferation of deep-learning-based technology, the potential risks associated with model development and…

Machine Learning · Computer Science 2023-01-06 Pin-Yu Chen , Sijia Liu