English
Related papers

Related papers: Behavioural Analysis of Alignment Faking

200 papers

Federated learning (FL) is a trending training paradigm to utilize decentralized training data. FL allows clients to update model parameters locally for several epochs, then share them to a global model for aggregation. This training…

Machine Learning · Computer Science 2022-08-09 Xiaoxiao Li , Zhao Song , Jiaming Yang

An important aspect in developing language models that interact with humans is aligning their behavior to be useful and unharmful for their human users. This is usually achieved by tuning the model in a way that enhances desired behaviors…

Computation and Language · Computer Science 2024-06-04 Yotam Wolf , Noam Wies , Oshri Avnery , Yoav Levine , Amnon Shashua

AI sycophancy is increasingly recognized as a harmful alignment, but research remains fragmented and underdeveloped at the conceptual level. This article redefines AI sycophancy as the tendency of large language models (LLMs) and other…

Human-Computer Interaction · Computer Science 2025-09-29 Lihua Du , Xing Lyu , Lezi Xie , Bo Feng

We find that correct-to-incorrect sycophancy signals are most linearly separable within multi-head attention activations. Motivated by the linear representation hypothesis, we train linear probes across the residual stream, multilayer…

Computation and Language · Computer Science 2026-01-26 Rifo Genadi , Munachiso Nwadike , Nurdaulet Mukhituly , Hilal Alquabeh , Tatsuya Hiraoka , Kentaro Inui

In multi-agent systems, agents need to interact and collaborate with other agents in environments. Agent modeling is crucial to facilitate agent interactions and make adaptive cooperation strategies. However, it is challenging for agents to…

Artificial Intelligence · Computer Science 2023-10-20 Baofu Fang , Caiming Zheng , Hao Wang

Language models increasingly appear to learn similar representations, despite differences in training objectives, architectures, and data modalities. This emerging compatibility between independently trained models introduces new…

Artificial Intelligence · Computer Science 2026-05-26 Matt Gorbett , Suman Jana

Recent work has developed optimization procedures to find token sequences, called adversarial triggers, which can elicit unsafe responses from aligned language models. These triggers are believed to be highly transferable, i.e., a trigger…

Computation and Language · Computer Science 2025-04-10 Nicholas Meade , Arkil Patel , Siva Reddy

Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that conflates helpfulness with polite submission. This latent bias, known as sycophancy, manifests as…

Computation and Language · Computer Science 2026-05-19 Sanskar Pandey , Ruhaan Chopra , Angkul Puniya , Sohom Pal

As AI adoption expands across human society, the problem of aligning AI models to match human preferences remains a grand challenge. Currently, the AI alignment field is deeply divided between behavioral and representational approaches,…

Computers and Society · Computer Science 2025-08-12 Ben Y. Reis , William La Cava

Deepfake detection remains a challenging task due to the difficulty of generalizing to new types of forgeries. This problem primarily stems from the overfitting of existing detection methods to forgery-irrelevant features and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zhiyuan Yan , Yong Zhang , Yanbo Fan , Baoyuan Wu

While the wide adoption of refusal training in large language models (LLMs) has showcased improvements in model safety, recent works have highlighted shortcomings due to the shallow nature of these alignment methods. To this end, the work…

Machine Learning · Computer Science 2026-04-17 Pankayaraj Pathmanathan , Furong Huang

Frontier AI agents may pursue hidden goals while concealing their pursuit from oversight. Alignment training aims to prevent such behavior by reinforcing the correct goals, but alignment may not always succeed and can lead to unwanted side…

Machine Learning · Computer Science 2026-02-27 Bruce W. Lee , Chen Yueh-Han , Tomek Korbak

Face anti-spoofing aims to discriminate the spoofing face images (e.g., printed photos) from live ones. However, adversarial examples greatly challenge its credibility, where adding some perturbation noise can easily change the predictions.…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Songlin Yang , Wei Wang , Chenye Xu , Ziwen He , Bo Peng , Jing Dong

Evaluating alignment in language models requires testing how they behave under realistic pressure, not just what they claim they would do. While alignment failures increasingly cause real-world harm, comprehensive evaluation frameworks with…

Artificial Intelligence · Computer Science 2026-02-25 Nora Petrova , John Burden

Processes are a crucial artefact in organizations, since they coordinate the execution of activities so that products and services are provided. The use of models to analyse the underlying processes is a well-known practice. However, due to…

Artificial Intelligence · Computer Science 2019-12-13 Thomas Chatain , Mathilde Boltenhagen , Josep Carmona

Prioritizing fairness is of central importance in artificial intelligence (AI) systems, especially for those societal applications, e.g., hiring systems should recommend applicants equally from different demographic groups, and risk…

Machine Learning · Computer Science 2022-03-04 Zhibo Wang , Xiaowei Dong , Henry Xue , Zhifei Zhang , Weifeng Chiu , Tao Wei , Kui Ren

Multimodal representation learning is fundamentally about transforming incomparable modalities into comparable representations. While prior research primarily focused on explicitly aligning these representations through targeted learning…

Machine Learning · Computer Science 2025-06-16 Megan Tjandrasuwita , Chanakya Ekbote , Liu Ziyin , Paul Pu Liang

Despite recent advances in Generative Adversarial Networks (GANs), with special focus to the Deepfake phenomenon there is no a clear understanding neither in terms of explainability nor of recognition of the involved models. In particular,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Luca Guarnera , Oliver Giudice , Matthias Niessner , Sebastiano Battiato

Federated learning has emerged as an umbrella term for centralized coordination strategies in multi-agent environments. While many federated learning architectures process data in an online manner, and are hence adaptive by nature, most…

Machine Learning · Computer Science 2020-05-06 Elsa Rizk , Stefan Vlaski , Ali H. Sayed

Chain-of-thought (CoT) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models faithfully externalizing their reasoning. Motivated by this vulnerability, we study…

Machine Learning · Computer Science 2026-05-18 Reilly Haskins , Bilal Chughtai , Joshua Engels
‹ Prev 1 3 4 5 6 7 10 Next ›