English
Related papers

Related papers: Behavioural Analysis of Alignment Faking

200 papers

As face recognition is widely used in diverse security-critical applications, the study of face anti-spoofing (FAS) has attracted more and more attention. Several FAS methods have achieved promising performances if the attack types in the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Yu-Chun Wang , Chien-Yi Wang , Shang-Hong Lai

In Federated Learning (FL), a group of workers participate to build a global model under the coordination of one node, the chief. Regarding the cybersecurity of FL, some attacks aim at injecting the fabricated local model updates into the…

Machine Learning · Computer Science 2021-11-30 Ranwa Al Mallah , Godwin Badu-Marfo , Bilal Farooq

Safety alignment is often conceptualized as a monolithic process wherein harmfulness detection automatically triggers refusal. However, the persistence of jailbreak attacks suggests a fundamental mechanistic decoupling. We propose the…

Cryptography and Security · Computer Science 2026-03-16 Jinman Wu , Yi Xie , Shen Lin , Shiqian Zhao , Xiaofeng Chen

Large language models often exhibit increased sycophantic behavior after preference-based post-training, showing a stronger tendency to affirm a user's stated or implied belief even when this conflicts with factual accuracy or sound…

Artificial Intelligence · Computer Science 2026-02-03 Itai Shapira , Gerdus Benade , Ariel D. Procaccia

Modern large language models rely on chain-of-thought (CoT) reasoning to achieve impressive performance, yet the same mechanism can amplify deceptive alignment, situations in which a model appears aligned while covertly pursuing misaligned…

Artificial Intelligence · Computer Science 2025-05-27 Jiaming Ji , Wenqi Chen , Kaile Wang , Donghai Hong , Sitong Fang , Boyuan Chen , Jiayi Zhou , Juntao Dai , Sirui Han , Yike Guo , Yaodong Yang

Distributed reinforcement learning policies face network delays, jitter, and packet loss when deployed across edge devices and cloud servers. Standard RL training assumes zero-latency interaction, causing severe performance degradation…

Machine Learning · Computer Science 2026-03-16 Carlos Purves , Pietro Lio'

Deepfake technology has raised concerns about the authenticity of digital content, necessitating the development of effective detection methods. However, the widespread availability of deepfakes has given rise to a new challenge in the form…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Sarwar Khan

Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet the field studies it without a shared foundation, conflating properties of the evaluation…

Machine Learning · Computer Science 2026-05-25 Changling Li , Terry Jingchen Zhang , Jie Zhang , Zhijing Jin , Sahar Abdelnabi , Maksym Andriushchenko

Despite the considerable performance improvements of face recognition algorithms in recent years, the same scientific advances responsible for this progress can also be used to create efficient ways to attack them, posing a threat to their…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Eduarda Caldeira , Guray Ozgur , Tahar Chettaoui , Marija Ivanovska , Peter Peer , Fadi Boutros , Vitomir Struc , Naser Damer

Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while valuable, raises concerns about sycophancy: the tendency to provide responses that…

Computation and Language · Computer Science 2026-04-14 Arya Shah , Deepali Mishra , Chaklam Silpasuwanchai

As Large Language Model (LLM) agents become more widespread, associated misalignment risks increase. While prior research has studied agents' ability to produce harmful outputs or follow malicious instructions, it remains unclear how likely…

Collective human movement is a hallmark of complex systems, exhibiting emergent order across diverse settings, from pedestrian flows to biological collectives. In high-speed scenarios, alignment interactions ensure efficient flow and…

Physics and Society · Physics 2025-06-03 Debasish Sarker , Yi Zhang , Lynn K. Perry , Daniel S. Messinger , Chaoming Song

Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Consistency training is a promising new alignment paradigm to mitigate such failures by…

Machine Learning · Computer Science 2026-05-22 Andy Han , Kristina Fujimoto , Avidan Shah , Kiet Nguyen , Kai Xu , Chen Yueh-Han , Ilia Sucholutsky , Rico Angell

Factorization-based models have gained popularity since the Netflix challenge {(2007)}. Since that, various factorization-based models have been developed and these models have been proven to be efficient in predicting users' ratings…

Artificial Intelligence · Computer Science 2024-05-15 Jinfeng Zhong , Elsa Negre

Preference-based alignment like Reinforcement Learning from Human Feedback (RLHF) learns from pairwise preferences, yet the labels are often noisy and inconsistent. Existing uncertainty-aware approaches weight preferences, but ignore a more…

Machine Learning · Computer Science 2026-01-27 Tiejin Chen , Xiaoou Liu , Vishnu Nandam , Kuan-Ru Liou , Hua Wei

Nowadays, the increasingly growing number of mobile and computing devices has led to a demand for safer user authentication systems. Face anti-spoofing is a measure towards this direction for bio-metric user authentication, and in…

Computer Vision and Pattern Recognition · Computer Science 2020-04-14 Suman Saha , Wenhao Xu , Menelaos Kanakis , Stamatios Georgoulis , Yuhua Chen , Danda Pani Paudel , Luc Van Gool

As Federated Learning (FL) expands to larger and more distributed environments, consistency in training is challenged by network-induced delays, clock unsynchronicity, and variability in client updates. This combination of factors may…

Machine Learning · Computer Science 2025-06-12 Baran Can Gül , Stefanos Tziampazis , Nasser Jazdi , Michael Weyrich

Federated Learning (FL) is a paradigm in Machine Learning (ML) that addresses data privacy, security, access rights and access to heterogeneous information issues by training a global model using distributed nodes. Despite its advantages,…

Cryptography and Security · Computer Science 2022-01-19 Ranwa Al Mallah , David Lopez , Godwin Badu Marfo , Bilal Farooq

Fine-tuning aligned language models on benign tasks unpredictably degrades safety guardrails, even when training data contains no harmful content and developers have no adversarial intent. We show that the prevailing explanation, that…

Diffusion models have emerged as powerful generative models in the text-to-image domain. This paper studies their application as observation-to-action models for imitating human behaviour in sequential environments. Human behaviour is…