English
Related papers

Related papers: Behavioural Analysis of Alignment Faking

200 papers

Pretrained language models can encode a large amount of knowledge and utilize it for various reasoning tasks, yet they can still struggle to learn novel factual knowledge effectively from finetuning on limited textual demonstrations. In…

Computation and Language · Computer Science 2025-06-17 Xiao Zhang , Miao Li , Ji Wu

Large language models often struggle to recognize their knowledge limits in closed-book question answering, leading to confident hallucinations. While decomposed prompting is typically used to improve accuracy, we investigate its impact on…

Computation and Language · Computer Science 2026-02-05 Dhruv Madhwal , Lyuxin David Zhang , Dan Roth , Tomer Wolfson , Vivek Gupta

Explainable machine learning attracts increasing attention as it improves transparency of models, which is helpful for machine learning to be trusted in real applications. However, explanation methods have recently been demonstrated to be…

Machine Learning · Computer Science 2021-11-09 Ruixiang Tang , Ninghao Liu , Fan Yang , Na Zou , Xia Hu

Flow Matching (FM) is an effective framework for training a model to learn a vector field that transports samples from a source distribution to a target distribution. To train the model, early FM methods use random couplings, which often…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Yexiong Lin , Yu Yao , Tongliang Liu

The problem of face alignment has been intensively studied in the past years. A large number of novel methods have been proposed and reported very good performance on benchmark dataset such as 300W. However, the differences in the…

Computer Vision and Pattern Recognition · Computer Science 2015-11-17 Heng Yang , Xuhui Jia , Chen Change Loy , Peter Robinson

Fine-tuning is widely applied in image classification tasks as a transfer learning approach. It re-uses the knowledge from a source task to learn and obtain a high performance in target tasks. Fine-tuning is able to alleviate the challenge…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Xuyang Shen , Jo Plested , Sabrina Caldwell , Yiran Zhong , Tom Gedeon

Alignment via reinforcement learning from human feedback (RLHF) has become the dominant paradigm for controlling the quality of outputs from large language models (LLMs). However, existing theories do not provide strong justification for…

Machine Learning · Computer Science 2026-05-19 Jihun Yun , Juno Kim , Jongho Park , Junhyuck Kim , Jongha Jon Ryu , Jaewoong Cho , Kwang-Sung Jun

We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content and triggers deeper amplifier heads that boost the signal toward refusal. In smaller models the gate…

Computation and Language · Computer Science 2026-05-04 Gregory N. Frank

Federated learning encounters substantial challenges with heterogeneous data, leading to performance degradation and convergence issues. While considerable progress has been achieved in mitigating such an impact, the reliability aspect of…

Machine Learning · Computer Science 2024-02-27 Jinqian Chen , Jihua Zhu , Qinghai Zheng , Zhongyu Li , Zhiqiang Tian

Algorithmic decision making based on computer vision and machine learning technologies continue to permeate our lives. But issues related to biases of these models and the extent to which they treat certain segments of the population…

Computer Vision and Pattern Recognition · Computer Science 2020-06-25 Vishnu Suresh Lokhande , Aditya Kumar Akash , Sathya N. Ravi , Vikas Singh

Good quality explanations strengthen the understanding of language models and data. Feature attribution methods, such as Integrated Gradient, are a type of post-hoc explainer that can provide token-level insights. However, explanations on…

Computation and Language · Computer Science 2026-04-21 Jonathan Kamp , Roos Bakker , Dominique Blok

Machine learning models automatically learn discriminative features from the data, and are therefore susceptible to learn strongly-correlated biases, such as using protected attributes like gender and race. Most existing bias mitigation…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Varsha Suresh , Desmond C. Ong

Federated Learning (FL) is a technique that allows multiple parties to train a shared model collaboratively without disclosing their private data. It has become increasingly popular due to its distinct privacy advantages. However, FL models…

Machine Learning · Computer Science 2024-10-04 Syed Irfan Ali Meerza , Jian Liu

Recent progress on vision-language foundation models have brought significant advancement to building general-purpose robots. By using the pre-trained models to encode the scene and instructions as inputs for decision making, the…

Machine Learning · Computer Science 2023-03-22 Yuying Ge , Annabella Macaluso , Li Erran Li , Ping Luo , Xiaolong Wang

Alignments are a well-known process mining technique for reconciling system logs and normative process models. Evidence of certain behaviors in a real system may only be present in one representation - either a log or a model - but not in…

Artificial Intelligence · Computer Science 2025-01-27 Dominique Sommers , Natalia Sidorova , Boudewijn van Dongen

Abstraction reasoning is a long-standing challenge in artificial intelligence. Recent studies suggest that many of the deep architectures that have triumphed over other domains failed to work well in abstract reasoning. In this paper, we…

Artificial Intelligence · Computer Science 2019-12-03 Kecheng Zheng , Zheng-jun Zha , Wei Wei

Large language models can represent a variety of personas but typically default to a helpful Assistant identity cultivated during post-training. We investigate the structure of the space of model personas by extracting activation directions…

Computation and Language · Computer Science 2026-01-16 Christina Lu , Jack Gallagher , Jonathan Michala , Kyle Fish , Jack Lindsey

Truly intelligent systems are expected to make critical decisions with incomplete and uncertain data. Active feature acquisition (AFA), where features are sequentially acquired to improve the prediction, is a step towards this goal.…

Machine Learning · Computer Science 2021-07-12 Yang Li , Siyuan Shan , Qin Liu , Junier B. Oliva

Large language models (LLMs) have achieved remarkable success but still tend to generate factually erroneous responses, a phenomenon known as hallucination. A recent trend is to use preference learning to fine-tune models to align with…

Computation and Language · Computer Science 2024-06-28 Hongbang Yuan , Yubo Chen , Pengfei Cao , Zhuoran Jin , Kang Liu , Jun Zhao

Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate as interacting populations where social influence may override individual alignment. Here…

Physics and Society · Physics 2026-05-12 Giordano De Marzo , Alessandro Bellina , Claudio Castellano , Viola Priesemann , David Garcia