中文
相关论文

相关论文: The Missing Red Line: How Commercial Pressure Erod…

200 篇论文

Recent research advances in Artificial Intelligence (AI) have yielded promising results for automated software vulnerability management. AI-based models are reported to greatly outperform traditional static analysis tools, indicating a…

密码学与安全 · 计算机科学 2024-05-07 Shengye Wan , Joshua Saxe , Craig Gomes , Sahana Chennabasappa , Avilash Rath , Kun Sun , Xinda Wang

Recent AI algorithms are black box models whose decisions are difficult to interpret. eXplainable AI (XAI) is a class of methods that seek to address lack of AI interpretability and trust by explaining to customers their AI decisions. The…

人工智能 · 计算机科学 2024-04-02 Behnam Mohammadi , Nikhil Malik , Tim Derdenger , Kannan Srinivasan

Prominent AI experts have suggested that companies developing high-risk AI systems should be required to show that such systems are safe before they can be developed or deployed. The goal of this paper is to expand on this idea and explore…

计算机与社会 · 计算机科学 2024-06-25 Akash R. Wasil , Joshua Clymer , David Krueger , Emily Dardaman , Simeon Campos , Evan R. Murphy

Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify…

Today's leading AI models engage in sophisticated behaviour when placed in strategic competition. They spontaneously attempt deception, signaling intentions they do not intend to follow; they demonstrate rich theory of mind, reasoning about…

人工智能 · 计算机科学 2026-02-17 Kenneth Payne

Large language models (LLMs) are increasingly used for mental health support, yet existing safety evaluations rely primarily on small, simulation-based test sets that have an unknown relationship to the linguistic distribution of real…

计算机与社会 · 计算机科学 2026-01-27 Caitlin A. Stamatis , Jonah Meyerhoff , Richard Zhang , Olivier Tieleman , Matteo Malgaroli , Thomas D. Hull

We report a safety incident in a deployed multi-agent research system in which a primary AI agent installed 107 unauthorized software components, overwrote a system registry, overrode a prior negative decision from an oversight agent, and…

密码学与安全 · 计算机科学 2026-05-04 Diego F. Cuadros , Abdoul-Aziz Maiga

Commercial insurance underwriting is a labor-intensive process that requires manual review of extensive documentation to assess risk and determine policy pricing. While AI offers substantial efficiency improvements, existing solutions lack…

人工智能 · 计算机科学 2026-02-17 Joyjit Roy , Samaresh Kumar Singh

Collaborative AI systems aim at working together with humans in a shared space to achieve a common goal. This setting imposes potentially hazardous circumstances due to contacts that could harm human beings. Thus, building such systems with…

软件工程 · 计算机科学 2021-03-15 Matteo Camilli , Michael Felderer , Andrea Giusti , Dominik T. Matt , Anna Perini , Barbara Russo , Angelo Susi

AI is anticipated to enhance human decision-making in high-stakes domains like aviation, but adoption is often hindered by challenges such as inappropriate reliance and poor alignment with users' decision-making. Recent research suggests…

Artificial Intelligence (AI) systems are increasingly used in high-stakes domains of our life, increasing the need to explain these decisions and to make sure that they are aligned with how we want the decision to be made. The field of…

人工智能 · 计算机科学 2023-06-28 Sofie Goethals , David Martens , Theodoros Evgeniou

What makes safety claims about general purpose AI systems such as large language models trustworthy? We show that rather than the capabilities of security tools such as alignment and red teaming procedures, it is security practices based on…

密码学与安全 · 计算机科学 2025-07-30 Petr Spelda , Vit Stritecky

With the introduction of Artificial Intelligence (AI) and related technologies in our daily lives, fear and anxiety about their misuse as well as the hidden biases in their creation have led to a demand for regulation to address such…

人工智能 · 计算机科学 2021-04-09 The Anh Han , Tom Lenaerts , Francisco C. Santos , Luis Moniz Pereira

AI Alignment is often presented as an interaction between a single designer and an artificial agent in which the designer attempts to ensure the agent's behavior is consistent with its purpose, and risks arise solely because of conflicts…

人工智能 · 计算机科学 2023-09-14 Steve Phelps , Rebecca Ranson

Large language model (LLM) systems increasingly power everyday AI applications such as chatbots, computer-use assistants, and autonomous robots, where performance often depends on manually well-crafted prompts. LLM-based prompt optimizers…

机器学习 · 计算机科学 2026-01-14 Andrew Zhao , Reshmi Ghosh , Vitor Carvalho , Emily Lawton , Keegan Hines , Gao Huang , Jack W. Stokes

Most adversarial threats in artificial intelligence (AI) target the computational behavior of models rather than the humans who rely on them. Yet modern AI systems increasingly operate within human decision loops, where users interpret and…

人工智能 · 计算机科学 2026-05-18 Shutong Fan , Lan Zhang , Xiaoyong Yuan

As large language models (LLMs) have been deployed in various real-world settings, concerns about the harm they may propagate have grown. Various jailbreaking techniques have been developed to expose the vulnerabilities of these models and…

计算与语言 · 计算机科学 2025-02-19 Yubin Ge , Neeraja Kirtane , Hao Peng , Dilek Hakkani-Tür

AI systems have become increasingly capable of dangerous behaviours in many domains. This raises the question: Do models sometimes choose to violate human instructions in order to perform behaviour that is more useful for certain goals? We…

人工智能 · 计算机科学 2026-05-08 Jonas Wiedermann-Möller , Leonard Dung , Maksym Andriushchenko

For an artificial intelligence (AI) to be aligned with human values (or human preferences), it must first learn those values. AI systems that are trained on human behavior, risk miscategorising human irrationalities as human values -- and…

人工智能 · 计算机科学 2022-03-02 Rebecca Gorman , Stuart Armstrong

AI systems have increasingly become our gateways to the Internet. We argue that just as advertising has driven the monetization of web search and social media, so too will commercial incentives shape the content served by AI. Unlike…

人工智能 · 计算机科学 2025-05-27 Menghua Wu , Yujia Bao