English
Related papers

Related papers: Honeypot Protocol

200 papers

Future AI deployments will likely be monitored for malicious behaviour. The ability of these AIs to subvert monitors by adversarially selecting against them - attack selection - is particularly concerning. To study this, we let a red team…

Cryptography and Security · Computer Science 2026-04-17 Joachim Schaeffer , Arjun Khandelwal , Tyler Tracy

As AI agents surpass human capabilities, scalable oversight -- the problem of effectively supplying human feedback to potentially superhuman AI models -- becomes increasingly critical to ensure alignment. While numerous scalable oversight…

Artificial Intelligence · Computer Science 2025-04-08 Abhimanyu Pallavi Sudhir , Jackson Kaunismaa , Arjun Panickssery

We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier AI company. We apply two complementary evaluations to four Claude models (Mythos…

Artificial Intelligence · Computer Science 2026-04-28 Robert Kirk , Alexandra Souly , Kai Fronsdal , Abby D'Cruz , Xander Davies

Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavior could be very hard to detect, and if present in future…

Machine Learning · Computer Science 2025-07-04 Mary Phuong , Roland S. Zimmermann , Ziyue Wang , David Lindner , Victoria Krakovna , Sarah Cogan , Allan Dafoe , Lewis Ho , Rohin Shah

Frontier models are increasingly trained and deployed as autonomous agent. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives - also known as scheming. We study whether…

Artificial Intelligence · Computer Science 2025-01-16 Alexander Meinke , Bronson Schoen , Jérémy Scheurer , Mikita Balesni , Rusheb Shah , Marius Hobbhahn

Windows operating systems (OS) are ubiquitous in enterprise Information Technology (IT) and operational technology (OT) environments. Due to their widespread adoption and known vulnerabilities, they are often the primary targets of malware…

Cryptography and Security · Computer Science 2025-05-02 Yan Lin Aung , Yee Loon Khoo , Davis Yang Zheng , Bryan Swee Duo , Sudipta Chattopadhyay , Jianying Zhou , Liming Lu , Weihan Goh

Contextual multi-armed bandits are classical models in reinforcement learning for sequential decision-making associated with individual information. A widely-used policy for bandits is Thompson Sampling, where samples from a data-driven…

Machine Learning · Statistics 2021-11-30 Hongju Park , Mohamad Kazem Shirani Faradonbeh

We study automated security response for an IT infrastructure and formulate the interaction between an attacker and a defender as a partially observed, non-stationary game. We relax the standard assumption that the game model is correctly…

Computer Science and Game Theory · Computer Science 2025-04-21 Kim Hammar , Tao Li , Rolf Stadler , Quanyan Zhu

Post-deployment monitoring of artificial intelligence (AI) systems in health care is essential to ensure their safety, quality, and sustained benefit-and to support governance decisions about which systems to update, modify, or…

The great performance of machine learning algorithms and deep neural networks in several perception and control tasks is pushing the industry to adopt such technologies in safety-critical applications, as autonomous robots and self-driving…

Machine Learning · Computer Science 2025-09-10 Giulio Rossolini , Alessandro Biondi , Giorgio Buttazzo

Lateral movement of advanced persistent threats has posed a severe security challenge. Due to the stealthy and persistent nature of the lateral movement, defenders need to consider time and spatial locations holistically to discover latent…

Networking and Internet Architecture · Computer Science 2020-10-07 Linan Huang , Quanyan Zhu

The ubiquitous nature of the IoT devices has brought serious security implications to its users. A lot of consumer IoT devices have little to no security implementation at all, thus risking user's privacy and making them target of mass…

Cryptography and Security · Computer Science 2018-12-14 Muhammad A. Hakim , Hidayet Aksu , A. Selcuk Uluagac , Kemal Akkaya

Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{evaluation awareness}. This raises concerns that models could…

Artificial Intelligence · Computer Science 2026-03-05 Maheep Chaudhary

We present a shared control paradigm that improves a user's ability to operate complex, dynamic systems in potentially dangerous environments without a priori knowledge of the user's objective. In this paradigm, the role of the autonomous…

Robotics · Computer Science 2019-06-07 Alexander Broad , Todd Murphey , Brenna Argall

A combination of deep reinforcement learning and supervised learning is proposed for the problem of active sequential hypothesis testing in completely unknown environments. We make no assumptions about the prior probability, the action and…

Artificial Intelligence · Computer Science 2023-06-07 George Stamatelis , Nicholas Kalouptsidis

The partial monitoring (PM) framework provides a theoretical formulation of sequential learning problems with incomplete feedback. On each round, a learning agent plays an action while the environment simultaneously chooses an outcome. The…

Machine Learning · Computer Science 2024-05-17 Maxime Heuillet , Ola Ahmad , Audrey Durand

Internet of Things (IoT) has seen a prolific rise in recent times and provides the ability to solve several key challenges faced by our societies and environment. Data produced by IoT provides a significant opportunity to infer context that…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-02 Ashish Manchanda , Prem Prakash Jayaraman , Abhik Banerjee , Arkady Zaslavsky , Shakthi Weerasinghe , Guang-Li Huang

The design of safe-critical control algorithms for systems under Denial-of-Service (DoS) attacks on the system output is studied in this work. We aim to address scenarios where attack-mitigation approaches are not feasible, and the system…

Systems and Control · Electrical Eng. & Systems 2023-11-14 Santiago Jimenez Leudo , Kunal Garg , Ricardo G. Sanfelice , Alvaro A. Cardenas

One of the widely used cyber deception techniques is decoying, where defenders create fictitious machines (i.e., honeypots) to lure attackers. Honeypots are deployed to entice attackers, but their effectiveness depends on their…

Cryptography and Security · Computer Science 2021-08-26 Palvi Aggarwal , Yinuo Du , Kuldeep Singh , Cleotilde Gonzalez

Honeypot is an important cyber defense technique that can expose attackers new attacks. However, the effectiveness of honeypots has not been systematically investigated, beyond the rule of thumb that their effectiveness depends on how they…

Cryptography and Security · Computer Science 2024-01-15 Md Mahabub Uz Zaman , Liangde Tao , Mark Maldonado , Chang Liu , Ahmed Sunny , Shouhuai Xu , Lin Chen
‹ Prev 1 4 5 6 7 8 10 Next ›