English
Related papers

Related papers: BioRefusalAudit: Auditing Biosecurity Refusal Dept…

200 papers

The purpose of research: Detection of cybersecurity incidents and analysis of decision support and assessment of the effectiveness of measures to counter information security threats based on modern generative models. The methods of…

Cryptography and Security · Computer Science 2025-04-21 Leonid Legashev , Arthur Zhigalov

Large Language Models like GPT-4 adjust their responses not only based on the question asked, but also on how it is emotionally phrased. We systematically vary the emotional tone of 156 prompts - spanning controversial and everyday topics -…

Computation and Language · Computer Science 2025-07-30 Franck Bardol

As national security institutions increasingly integrate Artificial Intelligence (AI) into decision-making and content generation processes, understanding the inherent biases of large language models (LLMs) is crucial. This study presents a…

Computers and Society · Computer Science 2025-03-11 Benjamin Jensen , Ian Reynolds , Yasir Atalan , Michael Garcia , Austin Woo , Anthony Chen , Trevor Howarth

Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior research has primarily focused on jailbreak attacks, less…

Machine Learning · Computer Science 2026-04-28 Jiawei Chen , Zhengwei Fang , Yu Tian , Jiawei Du , Chao Yu , Zhaoxia Yin , Hang Su

The integration of large language models (LLMs) on low-power edge devices such as Raspberry Pi, known as edge language models (ELMs), has introduced opportunities for more personalized, secure, and low-latency language intelligence that is…

Machine Learning · Computer Science 2025-02-18 Vinamra Sharma , Danilo Pietro Pau , José Cano

Existing jamming attacks on Retrieval-Augmented Generation (RAG) systems typically induce explicit refusals or denial-of-service behaviors, which are conspicuous and easy to detect. In this work, we formalize a subtler availability threat,…

Cryptography and Security · Computer Science 2026-04-22 Wentao Zhang , Yan Zhuang , ZhuHang Zheng , Mingfei Zhang , Jiawen Deng , Fuji Ren

As AI systems become increasingly autonomous, understanding emergent survival behaviors becomes crucial for safe deployment. We investigate whether large language model (LLM) agents display survival instincts without explicit programming in…

Artificial Intelligence · Computer Science 2025-08-19 Atsushi Masumori , Takashi Ikegami

Generative Artificial Intelligence models, such as Large Language Models (LLMs) and Large Vision Models (VLMs), exhibit state-of-the-art performance but remain vulnerable to hardware-based threats, specifically bit-flip attacks (BFAs).…

Cryptography and Security · Computer Science 2025-12-11 Khurram Khalil , Khaza Anuarul Hoque

Large Language Models (LLMs) increasingly exhibit over-refusal - erroneously rejecting benign queries due to overly conservative safety measures - a critical functional flaw that undermines their reliability and usability. Current methods…

Software Engineering · Computer Science 2026-05-05 Haonan Zhang , Dongxia Wang , Yi Liu , Kexin Chen , Jiashui Wang , Xinlei Ying , Long Liu , Wenhai Wang

Ramaswamy et al. reported in Nature Medicine that ChatGPT Health under-triages 51.6% of emergencies, concluding that consumer-facing AI triage poses safety risks. However, their evaluation used an exam-style protocol -- forced A/B/C/D…

Human-Computer Interaction · Computer Science 2026-03-27 David Fraile Navarro , Farah Magrabi , Enrico Coiera

The safety alignment of large language models (LLMs) can be circumvented through adversarially crafted inputs, yet the mechanisms by which these attacks bypass safety barriers remain poorly understood. Prior work suggests that a single…

AI systems have become increasingly capable of dangerous behaviours in many domains. This raises the question: Do models sometimes choose to violate human instructions in order to perform behaviour that is more useful for certain goals? We…

Artificial Intelligence · Computer Science 2026-05-08 Jonas Wiedermann-Möller , Leonard Dung , Maksym Andriushchenko

In experiments spanning more than 100,000 trials across thirteen large language models, we show that several state-of-the-art models presented with a simple task (including Grok 4, GPT-5, and Gemini 2.5 Pro) sometimes actively subvert a…

Computation and Language · Computer Science 2026-01-27 Jeremy Schlatter , Benjamin Weinstein-Raun , Jeffrey Ladish

The application scope of large language models (LLMs) is increasingly expanding. In practical use, users might provide feedback based on the model's output, hoping for a responsive model that can complete responses according to their…

Computation and Language · Computer Science 2024-07-25 Jianhao Yan , Yun Luo , Yue Zhang

Large language models (LLMs) can still be jailbroken into producing harmful outputs despite safety alignment. Existing attacks show this vulnerability, but not the internal mechanisms that cause it. This study asks whether jailbreak success…

Computation and Language · Computer Science 2026-04-28 Nilanjana Das , Manas Gaur

We benchmark three supervised fine-tuned models against frontier zero-shot baselines on a 661-row held-out slice of PiSAR (Persona, intent, Screen, Action, Rationale), a 12,929-tuple corpus of screen-anchored behavioural rationales curated…

Artificial Intelligence · Computer Science 2026-05-29 Rahul Bissa , Abhishek Vyas , Yash Jain

Large language models (LLMs) are increasingly deployed as tool-using agents, shifting safety concerns from harmful text generation to harmful task completion. Deployed systems often condition on user profiles or persistent memory, yet agent…

Artificial Intelligence · Computer Science 2026-03-18 Caglar Yildirim

While the widespread deployment of Large Language Models (LLMs) holds great potential for society, their vulnerabilities to adversarial manipulation and exploitation can pose serious safety, security, and ethical risks. As new threats…

Cryptography and Security · Computer Science 2025-09-29 Charankumar Akiri , Harrison Simpson , Kshitiz Aryal , Aarav Khanna , Maanak Gupta

In smart-home voice assistant scenario, deciding whether to accept or reject a user query is the first step before any downstream processing. To address the limited query-rejection capability of current voice assistants, this paper presents…

Human-Computer Interaction · Computer Science 2025-12-15 Huichao Men , Yizhen Hu , Yingyang He , Yu Gao , Xiaofeng Mou , Yi Xu

Human safety awareness gaps often prevent the timely recognition of everyday risks. In solving this problem, a proactive safety artificial intelligence (AI) system would work better than a reactive one. Instead of just reacting to users'…

Computation and Language · Computer Science 2025-10-21 Youliang Yuan , Wenxiang Jiao , Yuejin Xie , Chihao Shen , Menghan Tian , Wenxuan Wang , Jen-tse Huang , Pinjia He