中文
相关论文

相关论文: Evaluating whether AI models would sabotage AI saf…

200 篇论文

Modern AI-integrated IDEs are shifting from passive code completion to proactive Next Edit Suggestions (NES). Unlike traditional autocompletion, NES is designed to construct a richer context from both recent user interactions and the…

密码学与安全 · 计算机科学 2026-05-15 Yunlong Lyu , Yixuan Tang , Peng Chen , Tian Dong , Xinyu Wang , Zhiqiang Dong , Hao Chen

Large language models are being deployed as mental health support agents at scale, yet only 16% of LLM-based chatbot interventions have undergone rigorous clinical efficacy testing, and simulations reveal psychological deterioration in over…

计算与语言 · 计算机科学 2026-04-28 Suhas BN , Andrew M. Sherrill , Rosa I. Arriaga , Chris W. Wiese , Saeed Abdullah

A well-known limitation of AI systems is presumptuousness: the tendency of AI systems to provide confident answers when information may be lacking. This challenge is particularly acute in legal applications, where a core task for attorneys,…

人工智能 · 计算机科学 2026-04-23 Mohamed Afane , Emily Robitschek , Derek Ouyang , Daniel E. Ho

Large Language Model (LLM) safeguards, which implement request refusals, have become a widely adopted mitigation strategy against misuse. At the intersection of adversarial machine learning and AI safety, safeguard red teaming has…

密码学与安全 · 计算机科学 2025-06-10 Zifan Wang , Christina Q. Knight , Jeremy Kritz , Willow E. Primack , Julian Michael

Forecasting when AI systems will become capable of meaningfully accelerating AI research is a central challenge for AI safety. Existing benchmarks measure broad capability growth, but may not provide ample early warning signals for…

多智能体系统 · 计算机科学 2026-04-30 Joshua Sherwood , Ben Aybar , Benjamin Kaplan

The prospect of artificial intelligence (AI) competing in the adversarial landscape of cyber security has long been considered one of the most impactful, challenging, and potentially dangerous applications of AI. Here, we demonstrate a new…

System prompts for AI coding agents increasingly employ motivational framing -- from neutral task descriptions to fear-driven threats -- yet no controlled study has examined whether such framing affects agent behavior. We present two…

软件工程 · 计算机科学 2026-03-17 Wu Ji

As software development practices increasingly adopt AI-powered tools, ensuring that such tools can support secure coding has become critical. This study evaluates the effectiveness of GitHub Copilot's recently introduced code review…

软件工程 · 计算机科学 2025-09-18 Amena Amro , Manar H. Alalfi

This essay offers a philosophical analysis of the field of AI safety based on recent technical reports, with particular focus on Anthropic's study on "agentic misalignment" in frontier language models. It examines the recurring…

计算机与社会 · 计算机科学 2026-03-17 Mariana Lins Costa

As training artificial intelligence (AI) models is a lengthy and hence costly process, leakage of such a model's internal parameters is highly undesirable. In the case of AI accelerators, side-channel information leakage opens up the threat…

硬件体系结构 · 计算机科学 2024-12-11 Andrija Nešković , Saleh Mulhem , Alexander Treff , Rainer Buchty , Thomas Eisenbarth , Mladen Berekovic

Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe planning systematically, we introduce DESPITE, a benchmark of 12,279 tasks spanning physical…

人工智能 · 计算机科学 2026-05-05 Tao Zhang , Kaixian Qu , Zhibin Li , Jiajun Wu , Marco Hutter , Manling Li , Fan Shi

Current large language models (LLMs) excel in verifiable domains where outputs can be checked before action but prove less reliable for high-stakes strategic decisions with uncertain outcomes. This gap, driven by mutually reinforcing…

人工智能 · 计算机科学 2025-11-12 Alejandro R. Jadad

Are frontier AI systems becoming more capable? Certainly. Yet such progress is not an unalloyed blessing but rather a Trojan horse: behind their performance leaps lie more insidious and destructive safety risks, namely deception. Unlike…

人工智能 · 计算机科学 2026-05-28 Sitong Fang , Shiyi Hou , Kaile Wang , Boyuan Chen , Donghai Hong , Jiayi Zhou , Josef Dai , Yaodong Yang , Jiaming Ji

Chain-of-thought (CoT) reasoning has been proposed as a transparency mechanism for large language models in safety-critical deployments, yet its effectiveness depends on faithfulness (whether models accurately verbalize the factors that…

计算与语言 · 计算机科学 2026-03-25 Richard J. Young

Background: Due to their diversity, complexity, and above all importance, safety-critical and dependable systems must be developed with special diligence. Criticality increases as these systems likely contain artificial intelligence (AI)…

软件工程 · 计算机科学 2025-06-03 Amra Ramic , Stefan Kugele

As artificial intelligence (AI) systems become increasingly integral to organizational processes, they introduce new forms of fraud that are often subtle, systemic, and concealed within technical complexity. This paper introduces the…

计算机与社会 · 计算机科学 2025-08-20 Benjamin Zweers , Diptish Dey , Debarati Bhaumik

As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for…

人工智能 · 计算机科学 2025-05-12 Markov Grey , Charbel-Raphaël Segerie

Artificial Intelligence (AI) is revolutionizing scientific research, yet its growing integration into laboratory environments presents critical safety challenges. Large language models (LLMs) and vision language models (VLMs) now assist in…

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feedback to steer them towards desirable traits. One proposed…

人工智能 · 计算机科学 2025-05-26 Marie Davidsen Buhl , Jacob Pfau , Benjamin Hilton , Geoffrey Irving

AI is anticipated to enhance human decision-making in high-stakes domains like aviation, but adoption is often hindered by challenges such as inappropriate reliance and poor alignment with users' decision-making. Recent research suggests…