中文
相关论文

相关论文: Evaluating Frontier Models for Dangerous Capabilit…

200 篇论文

Ensuring responsible use of artificial intelligence (AI) has become imperative as autonomous systems increasingly influence critical societal domains. However, the concept of trustworthy AI remains broad and multi-faceted. This thesis…

人工智能 · 计算机科学 2025-10-28 Filip Cano

This chapter presents perspectives for challenges and future development in building reliable AI systems, particularly, agentic AI systems. Several open research problems related to mitigating the risks of cascading failures are discussed.…

人工智能 · 计算机科学 2025-11-18 Liudong Xing , Janet , Lin

Social engineering (SE) attacks remain a significant threat to both individuals and organizations. The advancement of Artificial Intelligence (AI), including diffusion models and large language models (LLMs), has potentially intensified…

密码学与安全 · 计算机科学 2024-07-24 Jingru Yu , Yi Yu , Xuhong Wang , Yilun Lin , Manzhi Yang , Yu Qiao , Fei-Yue Wang

This chapter starts with a sketch of how we got to "generative AI" (GenAI) and a brief summary of the various impacts it had so far. It then discusses some of the opportunities of GenAI, followed by the challenges and dangers, including…

计算机与社会 · 计算机科学 2025-08-19 Matthias Scheutz

There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluation ecosystem is insufficient: Commonly used static benchmarks…

The conversation around artificial intelligence (AI) often focuses on safety, transparency, accountability, alignment, and responsibility. However, AI security (i.e., the safeguarding of data, models, and pipelines from adversarial…

密码学与安全 · 计算机科学 2025-04-24 Krti Tallam

AI systems have become increasingly capable of dangerous behaviours in many domains. This raises the question: Do models sometimes choose to violate human instructions in order to perform behaviour that is more useful for certain goals? We…

人工智能 · 计算机科学 2026-05-08 Jonas Wiedermann-Möller , Leonard Dung , Maksym Andriushchenko

The recent development of powerful AI systems has highlighted the need for robust risk management frameworks in the AI industry. Although companies have begun to implement safety frameworks, current approaches often lack the systematic…

人工智能 · 计算机科学 2025-02-20 Simeon Campos , Henry Papadatos , Fabien Roger , Chloé Touzet , Otter Quarks , Malcolm Murray

We develop a taxonomical framework for classifying challenges to the possibility of consciousness in digital artificial intelligence systems. This framework allows us to identify the level of granularity at which a given challenge is…

人工智能 · 计算机科学 2025-11-21 Andres Campero , Derek Shiller , Jaan Aru , Jonathan Simon

The dawn of Generative Artificial Intelligence (GAI), characterized by advanced models such as Generative Pre-trained Transformers (GPT) and other Large Language Models (LLMs), has been pivotal in reshaping the field of data analysis,…

密码学与安全 · 计算机科学 2024-05-06 Shivani Metta , Isaac Chang , Jack Parker , Michael P. Roman , Arturo F. Ehuan

The widespread utilization of AI systems has drawn attention to the potential impacts of such systems on society. Of particular concern are the consequences that prediction errors may have on real-world scenarios, and the trust humanity…

计算机与社会 · 计算机科学 2021-06-22 Mary Roszel , Robert Norvill , Jean Hilger , Radu State

Risk assessments for advanced AI systems require evaluating both the models themselves and their deployment contexts. We introduce the Societal Capacity Assessment Framework (SCAF), an indicators-based approach to measuring a society's…

计算机与社会 · 计算机科学 2025-09-30 Milan Gandhi , Peter Cihon , Owen Larter , Rebecca Anselmetti

AI Safety is an emerging area of critical importance to the safe adoption and deployment of AI systems. With the rapid proliferation of AI and especially with the recent advancement of Generative AI (or GAI), the technology ecosystem behind…

人工智能 · 计算机科学 2026-05-14 Chen Chen , Xueluan Gong , Ziyao Liu , Weifeng Jiang , Si Qi Goh , Kwok-Yan Lam

The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing remains nascent and is still a developing science. As AI agents…

Robots applications in our daily life increase at an unprecedented pace. As robots will soon operate "out in the wild", we must identify the safety and security vulnerabilities they will face. Robotics researchers and manufacturers focus…

机器人学 · 计算机科学 2021-11-18 Michele Colledanchise

There are many goals for an AI that could become dangerous if the AI becomes superintelligent or otherwise powerful. Much work on the AI control problem has been focused on constructing AI goals that are safe even for such AIs. This paper…

人工智能 · 计算机科学 2017-05-31 Stuart Armstrong , Benjamin Levinstein

AI safety benchmarks are pivotal for safety in advanced AI systems; however, they have significant technical, epistemic, and sociotechnical shortcomings. We present a review of 210 safety benchmarks that maps out common challenges in safety…

计算机与社会 · 计算机科学 2026-02-10 Cheng Yu , Severin Engelmann , Ruoxuan Cao , Dalia Ali , Orestis Papakyriakopoulos

In robotics, one of the main challenges is that the on-board Artificial Intelligence (AI) must deal with different or unexpected environments. Such AI agents may be incompetent there, while the underlying model itself may not be aware of…

机器人学 · 计算机科学 2020-05-05 Gertjan J. Burghouts , Albert Huizing , Mark A. Neerincx

Mitigating the risks from frontier AI systems requires up-to-date and reliable information about those systems. Organizations that develop and deploy frontier systems have significant access to such information. By reporting safety-critical…