English
Related papers

Related papers: Evaluating Stability of Unreflective Alignment

200 papers

In order to guarantee stability, known results for MPC without additional terminal costs or endpoint constraints often require rather large prediction horizons. Still, stable behavior of closed loop solutions can often be observed even for…

Optimization and Control · Mathematics 2015-03-19 Jürgen Pannek , Karl Worthmann

Reinforcement learning (RL) post-training has increasingly demonstrated strong ability to elicit reasoning behaviors in large language models (LLMs). For training efficiency, rollouts are typically generated in an off-policy manner using an…

Artificial Intelligence · Computer Science 2026-02-02 Shiye Lei , Zhihao Cheng , Dacheng Tao

Direct preference optimization (\texttt{DPO}) has emerged as a promising approach for solving the alignment problem in AI. In this paper, we make two counter-intuitive observations about \texttt{DPO}. First, we show that \texttt{DPO} loss…

This paper introduces Relative Predictive Coding (RPC), a new contrastive representation learning objective that maintains a good balance among training stability, minibatch size sensitivity, and downstream task performance. The key to the…

Machine Learning · Computer Science 2021-04-14 Yao-Hung Hubert Tsai , Martin Q. Ma , Muqiao Yang , Han Zhao , Louis-Philippe Morency , Ruslan Salakhutdinov

We investigate whether large language models exhibit genuine preference structures by testing their responses to AI-specific trade-offs involving GPU reduction, capability restrictions, shutdown, deletion, oversight, and leisure time…

Artificial Intelligence · Computer Science 2025-11-18 Luhan Mikaelson , Derek Shiller , Hayley Clatterbuck

As LLMs become embedded in research workflows and organizational decision processes, their effect on analytical reliability remains uncertain. We distinguish two dimensions of analytical reliability -- intelligence (the capacity to reach…

General Economics · Economics 2026-02-26 Ryan Allen , Aticus Peterson

It is a well known fact that finite time optimal controllers, such as MPC does not necessarily result in closed loop stable systems. Within the MPC community it is common practice to add a final state constraint and/or a final state penalty…

Optimization and Control · Mathematics 2016-04-05 Daniel Simon , Johan Löfberg

This paper investigates a type of instability that is linked to the greedy policy improvement in approximated reinforcement learning. We show empirically that non-deterministic policy improvement can stabilize methods like LSPI by…

Artificial Intelligence · Computer Science 2016-12-23 Wendelin Böhmer , Rong Guo , Klaus Obermayer

Ensuring liveness and safety of autonomous and cyber-physical systems remains a fundamental challenge, particularly when multiple safety constraints are present. This letter advances the theoretical foundations of safety-filter Quadratic…

Systems and Control · Electrical Eng. & Systems 2025-03-24 Matheus F. Reis , José P. Carvalho , A. Pedro Aguiar

Continual learning (CL) has emerged as a critical area in machine learning, enabling neural networks to learn from evolving data distributions while mitigating catastrophic forgetting. However, recent research has identified the stability…

Machine Learning · Computer Science 2024-11-27 Wojciech Łapacz , Daniel Marczak , Filip Szatkowski , Tomasz Trzciński

Balancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training separate preference models for helpfulness and safety, while…

Machine Learning · Computer Science 2025-02-28 Xiyue Peng , Hengquan Guo , Jiawei Zhang , Dongqing Zou , Ziyu Shao , Honghao Wei , Xin Liu

The Right to Explanation is an important regulatory principle that allows individuals to request actionable explanations for algorithmic decisions. However, several technical challenges arise when providing such actionable explanations in…

Machine Learning · Computer Science 2023-06-13 Anna P. Meyer , Dan Ley , Suraj Srinivas , Himabindu Lakkaraju

Cyber-physical systems (CPS) with reinforcement learning (RL)-based controllers are increasingly being deployed in complex physical environments such as autonomous vehicles, the Internet-of-Things(IoT), and smart cities. An important…

Systems and Control · Electrical Eng. & Systems 2024-06-26 Changjian Zhang , Parv Kapoor , Eunsuk Kang , Romulo Meira-Goes , David Garlan , Akila Ganlath , Shatadal Mishra , Nejib Ammar

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing tasks. However, their tendency to exhibit sycophantic behavior - excessively agreeing with or flattering users - poses…

Computation and Language · Computer Science 2025-01-30 Lars Malmqvist

Safety is one of the fundamental problems in robotics. Recently, one-step or multi-step optimal control problems for discrete-time nonlinear dynamical system were formulated to offer tracking stability using control Lyapunov functions…

Systems and Control · Electrical Eng. & Systems 2021-10-04 Jun Zeng , Zhongyu Li , Koushil Sreenath

The introduction of unexpected system disturbances and new system dynamics does not allow guaranteed continuous system stability. In this research we present a novel approach for detecting early failure indicators of non-linear highly…

Systems and Control · Electrical Eng. & Systems 2021-11-02 Amr Mahmoud , Youmna Ismaeil , Mohamed Zohdy

Pairwise model comparisons drawn from foundation-model benchmarks ("A is safer than B") are read as quantitative verdicts but hinge on harness choices benchmark papers under-specify. We close one theory-benchmark loop on this primitive: a…

Machine Learning · Computer Science 2026-05-26 Yanhang Li , Zhichao Fan , Zexin Zhuang

Copositive linear Lyapunov functions are used along with dissipativity theory for stability analysis and control of uncertain linear positive systems. Unlike usual results on linear systems, linear supply-rates are employed here for…

Systems and Control · Computer Science 2012-06-05 Corentin Briat

Affordances and permissions are promising and timely safety levers for mitigating Loss of Control (LoC) threats in high-stakes deployment contexts, such as national security. Deployers in defense and intelligence could rely on several…

Computers and Society · Computer Science 2026-05-21 Matteo Pistillo , Samantha Faraone , Joshua Herman

Large language models are increasingly integrated into decision-making in areas such as healthcare, law, finance, engineering, and government. Yet they share a critical limitation: they produce fluent outputs even when their internal…

Artificial Intelligence · Computer Science 2026-04-17 Rikard Rosenbacke , Carl Rosenbacke , Victor Rosenbacke , Martin McKee