中文
相关论文

相关论文: "Think First, Verify Always": Training Humans to F…

200 篇论文

This document focuses on the threats, especially near-term threats, that Artificial Intelligence (AI) brings to society. Most of the threats discussed here can result from any algorithmic process, not just AI; in addition, defining AI is…

计算机与社会 · 计算机科学 2024-09-10 Don Byrd

For an artificial intelligence (AI) to be aligned with human values (or human preferences), it must first learn those values. AI systems that are trained on human behavior, risk miscategorising human irrationalities as human values -- and…

人工智能 · 计算机科学 2022-03-02 Rebecca Gorman , Stuart Armstrong

Computer security has been a concern for decades and artificial intelligence techniques have been applied to the area for nearly as long. Most of the techniques are being applied to the detection of attacks to running systems, but recent…

密码学与安全 · 计算机科学 2019-12-17 Steve Kommrusch

Users' physical safety is an increasing concern as the market for intelligent systems continues to grow, where unconstrained systems may recommend users dangerous actions that can lead to serious injury. Covertly unsafe text is an area of…

计算与语言 · 计算机科学 2023-05-22 Alex Mei , Sharon Levy , William Yang Wang

In this position paper, I argue that the best way to help and protect humans using AI technology is to make them aware of the intrinsic limitations and problems of AI algorithms. To accomplish this, I suggest three ethical guidelines to be…

计算机与社会 · 计算机科学 2021-12-03 Claudio S. Pinhanez

AI Safety has become a vital front-line concern of many scientists within and outside the AI community. There are many immediate and long term anticipated risks that range from existential risk to human existence to deep fakes and bias in…

人工智能 · 计算机科学 2024-10-15 Simon Kasif

Generative Artificial Intelligence (GenAI) presents significant advancements but also introduces novel security challenges, particularly within agentic workflows where AI agents operate autonomously. These risks escalate in multi-agent…

密码学与安全 · 计算机科学 2025-06-24 Sunil Kumar Jang Bahadur , Gopala Dhar

Research in cybersecurity may seem reactive, specific, ephemeral, and indeed ineffective. Despite decades of innovation in defense, even the most critical software systems turn out to be vulnerable to attacks. Time and again. Offense and…

密码学与安全 · 计算机科学 2024-09-04 Marcel Böhme

Frontier AI systems are rapidly advancing in their capabilities to persuade, deceive, and influence human behaviour, with current models already demonstrating human-level persuasion and strategic deception in specific contexts. Humans are…

This second update to the 2025 International AI Safety Report assesses new developments in general-purpose AI risk management over the past year. It examines how researchers, public institutions, and AI developers are approaching risk…

In many contexts, lying -- the use of verbal falsehoods to deceive -- is harmful. While lying has traditionally been a human affair, AI systems that make sophisticated verbal statements are becoming increasingly prevalent. This raises the…

计算机与社会 · 计算机科学 2021-10-14 Owain Evans , Owen Cotton-Barratt , Lukas Finnveden , Adam Bales , Avital Balwit , Peter Wills , Luca Righetti , William Saunders

The emergence of AI tools in cybersecurity creates many opportunities and uncertainties. A focus group with advanced graduate students in cybersecurity revealed the potential depth and breadth of the challenges and opportunities. The…

密码学与安全 · 计算机科学 2023-11-03 Diane Jackson , Sorin Adam Matei , Elisa Bertino

The history of AI has included several "waves" of ideas. The first wave, from the mid-1950s to the 1980s, focused on logic and symbolic hand-encoded representations of knowledge, the foundations of so-called "expert systems". The second…

计算机与社会 · 计算机科学 2020-12-14 Odest Chadwicke Jenkins , Daniel Lopresti , Melanie Mitchell

Keeping up with threat intelligence is a must for a security analyst today. There is a volume of information present in `the wild' that affects an organization. We need to develop an artificial intelligence system that scours the…

人工智能 · 计算机科学 2019-05-09 Sudip Mittal , Anupam Joshi , Tim Finin

This paper presents a novel, structured decision support framework that systematically aligns diverse artificial intelligence (AI) agent architectures, reactive, cognitive, hybrid, and learning, with the comprehensive National Institute of…

人工智能 · 计算机科学 2025-10-03 Masike Malatji

As artificial intelligence (AI) capabilities advance rapidly, frontier models increasingly demonstrate systematic deception and scheming, complying with safety protocols during oversight but defecting when unsupervised. This paper examines…

计算机与社会 · 计算机科学 2026-02-20 Josef A. Habdank

In today's increasingly digital interactions, robust Identity Verification (IDV) is crucial for security and trust. Artificial Intelligence (AI) is transforming IDV, enhancing accuracy and fraud detection. This paper introduces ``Zero to…

密码学与安全 · 计算机科学 2025-03-13 Aniket Vaidya , Anurag Awasthi

Increasingly sophisticated and varied cyber threats necessitate ever improving enterprise security postures. For many organizations today, those postures have a foundation in the Zero Trust Architecture. This strategy sees trust as…

密码学与安全 · 计算机科学 2025-08-19 Samuel Aiello

Well-designed technologies that offer high levels of human control and high levels of computer automation can increase human performance, leading to wider adoption. The Human-Centered Artificial Intelligence (HCAI) framework clarifies how…

人机交互 · 计算机科学 2020-02-25 Ben Shneiderman