English
Related papers

Related papers: Behavioral Fingerprints for LLM Endpoint Stability…

200 papers

AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing has improved the behavior of modern language models, but aligned behavior does not by…

Artificial Intelligence · Computer Science 2026-05-27 Yige Li , Yunhao Feng , Jun Sun

The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a proliferation of safety benchmarks since late 2023. However, these benchmarks have developed…

Computers and Society · Computer Science 2026-05-19 Miles Q. Li , Benjamin C. M. Fung , Boyang Li , Heba Ismail , Farkhund Iqbal

Many industrial and security applications employ a suite of sensors for detecting abrupt changes in temporal behavior patterns. These abrupt changes typically manifest locally, rendering only a small subset of sensors informative.…

Machine Learning · Computer Science 2023-06-14 Aditya Gopalan , Venkatesh Saligrama , Braghadeesh Lakshminarayanan

AI models of equivalent capability can exhibit fundamentally different behavioral patterns, yet no standardized instrument exists to measure these dispositional differences. Existing approaches either borrow human personality dimensions and…

Artificial Intelligence · Computer Science 2026-04-03 Jihoon Jeong

Learned optimizers -- neural networks that are trained to act as optimizers -- have the potential to dramatically accelerate training of machine learning models. However, even when meta-trained across thousands of tasks at huge…

Machine Learning · Computer Science 2022-09-23 James Harrison , Luke Metz , Jascha Sohl-Dickstein

In many computational tasks and dynamical systems, asynchrony and randomization are naturally present and have been considered as ways to increase the speed and reduce the cost of computation while compromising the accuracy and convergence…

Machine Learning · Computer Science 2020-12-09 Sahin Lale , Oguzhan Teke , Babak Hassibi , Anima Anandkumar

In this paper, we study the simultaneous stability problem of a finite number of locally inter-connected linear subsystems under practical constraints, including asynchronous and aperiodic sampling, time-varying delays, and measurement…

Optimization and Control · Mathematics 2017-10-06 Feng Xiao , Yang Shi , Wei Ren

We provide a thorough study of stability of the 1-D continuity equation, which models many physical conservation laws. In our system-theoretic perspective, the velocity is considered to be an input. An additional input appears in the…

Optimization and Control · Mathematics 2019-08-19 Iasson Karafyllis , Miroslav Krstic

Understanding simplicity biases in deep learning offers a promising path toward developing reliable AI. A common metric for this, inspired by Boolean function analysis, is average sensitivity, which captures a model's robustness to…

Machine Learning · Computer Science 2026-02-10 Themistoklis Haris , Zihan Zhang , Yuichi Yoshida

Typical evaluations of fingerprint recognition systems consist of end-to-end black-box evaluations, which assess performance in terms of overall identification or authentication accuracy. However, these black-box tests of system performance…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Steven A. Grosz , Joshua J. Engelsma , Anil K. Jain

Finding the best way of adapting pre-trained language models to a task is a big challenge in current NLP. Just like the previous generation of task-tuned models (TT), models that are adapted to tasks via in-context-learning (ICL) are robust…

Computation and Language · Computer Science 2023-10-23 Lucas Weber , Elia Bruni , Dieuwke Hupkes

Dynamic systems in AI are often complex and heterogeneous, so that an internal specification is not accessible and verification techniques such as model checking are not applicable. Monitoring is in such cases an attractive alternative, as…

Artificial Intelligence · Computer Science 2026-05-15 Alessandro Gianola , Marco Montali , Sarah Winkler

Behavioral evaluation is the dominant paradigm for assessing alignment in large language models (LLMs). In current practice, observed compliance under finite evaluation protocols is treated as evidence of latent alignment. However, the…

Machine Learning · Computer Science 2026-02-10 Igor Santos-Grueiro

This contribution is a follow-up of a recent paper by the authors on adaptive, non-linear time-frequency transforms, focusing on the STFT based transforms. The adaptivity is provided by a focus function, that depends on the analyzed…

Classical Analysis and ODEs · Mathematics 2025-06-11 Pierre Warion , Bruno Torrésani

Predictive Business Process Monitoring (PBPM) aims to forecast future outcomes of ongoing business processes. However, existing methods often lack flexibility to handle real-world challenges such as simultaneous events, class imbalance, and…

Machine Learning · Computer Science 2025-08-06 Fang Wang , Paolo Ceravolo , Ernesto Damiani

Ensuring the long-term reproducibility of data analyses requires results stability tests to verify that analysis results remain within acceptable variation bounds despite inevitable software updates and hardware evolutions. This paper…

Many theoretical obstacles to AI alignment are consequences of reflective stability - the problem of designing alignment mechanisms that the AI would not disable if given the option. However, problems stemming from reflective stability are…

Artificial Intelligence · Computer Science 2024-08-28 James Lucassen , Mark Henry , Philippa Wright , Owen Yeung

In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. However, they largely fail to assess whether the output image…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Jiayang Li , Shuo Cao , Xiaohui Li , Zhizhen Zhang , Kaiwen Zhu , Yule Duan , Yu Qiao , Jian Zhang , Yihao Liu

This study introduces the Kerimov-Alekberli model, a novel information-geometric framework that redefines AI safety by formally linking non-equilibrium thermodynamics to stochastic control for the ethical alignment of autonomous systems. By…

Artificial Intelligence · Computer Science 2026-05-06 Hikmat Karimov , Rahid Zahid Alekberli

In mechanistic interpretability, recent work scrutinizes transformer "circuits" - sparse, mono or multi layer sub computations, that may reflect human understandable functions. Yet, these network circuits are rarely acid-tested for their…

Machine Learning · Computer Science 2026-02-20 Karan Bali , Jack Stanley , Praneet Suresh , Danilo Bzdok