English
Related papers

Related papers: Stress-Testing Alignment Audits With Prompt-Level …

200 papers

Automated Driving Systems (ADS), including Advanced Driver Assistance Systems (ADAS), must fulfill not only high functional expectations but also stringent timing constraints mandated by international regulations and standards. Regulatory…

Software Engineering · Computer Science 2026-05-05 Sebastian Dingler , Philip Rehkop , Florian Mayer , Ralf Muenzenberger

We investigate strategic deception in large language models using two complementary testbeds: Secret Agenda (across 38 models) and Insider Trading compliance (via SAE architectures). Secret Agenda reliably induced lying when deception…

Computers and Society · Computer Science 2025-09-26 Caleb DeLeeuw , Gaurav Chawla , Aniket Sharma , Vanessa Dietze

TRUST Agents is a collaborative multi-agent framework for explainable fact verification and fake news detection. Rather than treating verification as a simple true-or-false classification task, the system identifies verifiable claims,…

Artificial Intelligence · Computer Science 2026-04-15 Gautama Shastry Bulusu Venkata , Santhosh Kakarla , Maheedhar Omtri Mohan , Aishwarya Gaddam

LLM-based financial agents increasingly produce investment rationales before the outcomes needed to evaluate them are observable. This creates a delayed-ground-truth evaluation problem: realized returns remain the eventual arbiter of…

Artificial Intelligence · Computer Science 2026-05-05 Sidi Chang , Peiying Zhu , Yuxiao Chen

Financial risk detection in Enterprise Resource Planning (ERP) systems is an important but underexplored application of machine learning. Published studies in this area tend to suffer from vague dataset descriptions, leakage-prone…

Machine Learning · Computer Science 2026-03-10 Sanjay Mishra

Adversarial robustness is one of the essential safety criteria for guaranteeing the reliability of machine learning models. While various adversarial robustness testing approaches were introduced in the last decade, we note that most of…

Machine Learning · Statistics 2022-04-04 Giuseppe Castiglione , Gavin Ding , Masoud Hashemi , Christopher Srinivasa , Ga Wu

Auditing fairness of decision-makers is now in high demand. To respond to this social demand, several fairness auditing tools have been developed. The focus of this study is to raise an awareness of the risk of malicious decision-makers who…

Machine Learning · Statistics 2019-12-02 Kazuto Fukuchi , Satoshi Hara , Takanori Maehara

We stress-tested 16 leading models from multiple developers in hypothetical corporate environments to identify potentially risky agentic behaviors before they cause real harm. In the scenarios, we allowed models to autonomously send emails…

Cryptography and Security · Computer Science 2025-10-17 Aengus Lynch , Benjamin Wright , Caleb Larson , Stuart J. Ritchie , Soren Mindermann , Evan Hubinger , Ethan Perez , Kevin Troy

LLM-based financial agents increasingly rely on both numerical market data and textual signals for sequential trading and stock prediction. However, financial misinformation often appears as subtle textual perturbations rather than explicit…

Computational Engineering, Finance, and Science · Computer Science 2026-05-12 Zhiwei Liu , Yangyang Yu , Yupeng Cao , Yuechen Jiang , Haohang Li , Zhuoran Lu , Yuyan Wang , Yixiang Zheng , Xiaorui Guo , Calvin Yixiang Cheng , Sophia Ananiadou

Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret collusion whenever doing so confers a strategic advantage. To investigate this phenomenon, we…

Artificial Intelligence · Computer Science 2026-05-28 Xijie Zeng , Frank Rudzicz

Despite significant advances in alignment techniques, we demonstrate that state-of-the-art language models remain vulnerable to carefully crafted conversational scenarios that can induce various forms of misalignment without explicit…

Computation and Language · Computer Science 2025-08-07 Siddhant Panpatil , Hiskias Dingeto , Haon Park

As Large Language Models transition to autonomous agents, user inputs frequently violate cooperative assumptions (e.g., implicit intent, missing parameters, false presuppositions, or ambiguous expressions), creating execution risks that…

Artificial Intelligence · Computer Science 2026-02-03 Han Bao , Zheyuan Zhang , Pengcheng Jing , Zhengqing Yuan , Kaiwen Shi , Yanfang Ye

Monumental advancements in artificial intelligence (AI) have lured the interest of doctors, lenders, judges, and other professionals. While these high-stakes decision-makers are optimistic about the technology, those familiar with AI…

Artificial Intelligence · Computer Science 2023-04-13 Zachariah Carmichael , Walter J Scheirer

Activation-based probes have emerged as a promising approach for detecting deceptively aligned AI systems by identifying internal conflict between true and stated goals. We identify a fundamental blind spot: probes fail on coherent…

Machine Learning · Computer Science 2026-03-30 Kristiyan Haralambiev

LLM-based agents increasingly coordinate decisions in multi-agent systems, often attaching natural-language reasoning to actions. However, reasoning is neither free nor automatically reliable: it incurs computational cost and, without…

Multiagent Systems · Computer Science 2026-04-14 Feliks Bańka , Jarosław A. Chudziak

Larger language models (LLMs) have taken the world by storm with their massive multi-tasking capabilities simply by optimizing over a next-word prediction objective. With the emergence of their properties and encoded knowledge, the risk of…

Computation and Language · Computer Science 2023-08-31 Rishabh Bhardwaj , Soujanya Poria

Empowering large language models to accurately express confidence in their answers is essential for trustworthy decision-making. Previous confidence elicitation methods, which primarily rely on white-box access to internal model information…

Computation and Language · Computer Science 2024-03-19 Miao Xiong , Zhiyuan Hu , Xinyang Lu , Yifei Li , Jie Fu , Junxian He , Bryan Hooi

The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Shared pretraining data, distillation, and alignment pipelines can induce hidden behavioral…

Artificial Intelligence · Computer Science 2026-04-10 Chenchen Kuai , Jiwan Jiang , Zihao Zhu , Hao Wang , Keshu Wu , Zihao Li , Yunlong Zhang , Chenxi Liu , Zhengzhong Tu , Zhiwen Fan , Yang Zhou

With the widespread application of LLM-based agents across various domains, their complexity has introduced new security threats. Existing red-team methods mostly rely on modifying user prompts, which lack adaptability to new data and may…

Computation and Language · Computer Science 2026-04-14 Yanxu Mao , Peipei Liu , Tiehan Cui , Congying Liu , Mingzhe Xing , Datao You

Systems operating in adversarial environments may inadvertently leak sensitive information to adversaries. To address this challenge, we revisit the linear-quadratic control framework and introduce deception to actively mislead adversaries.…

Optimization and Control · Mathematics 2026-04-02 Yerin Kim , Haosheng Zhou , Alexander Benvenuti , Ruimeng Hu , Matthew Hale
‹ Prev 1 3 4 5 6 7 10 Next ›