English
Related papers

Related papers: Robustness to fundamental uncertainty in AGI align…

200 papers

Self-consistent approaches in many-electron problems typically converge to an unphysical solution in strongly correlated regimes. By deriving the mathematical condition for the stability of the physical solution, we unveil the precise…

Strongly Correlated Electrons · Physics 2026-04-27 Herbert Eßl , Matthias Reitner , Evgeny Kozik , Alessandro Toschi

This work proposes a framework for multistage adjustable robust optimization that unifies the treatment of three different types of endogenous uncertainty, where decisions, respectively, (i) alter the uncertainty set, (ii) affect the…

Optimization and Control · Mathematics 2020-08-31 Qi Zhang , Wei Feng

General intelligence, the ability to solve arbitrary solvable problems, is supposed by many to be artificially constructible. Narrow intelligence, the ability to solve a given particularly difficult problem, has seen impressive recent…

Artificial Intelligence · Computer Science 2020-07-22 Michael K Cohen , Badri Vellambi , Marcus Hutter

Recent approaches to evaluating Artificial General Intelligence (AGI) typically summarize a system's capability using the arithmetic mean of its proficiencies across multiple cognitive domains. While simple, this implicitly assumes…

Artificial Intelligence · Computer Science 2025-12-01 Fares Fourati

As artificial intelligence systems move toward clinical deployment, ensuring reliable prediction behavior is fundamental for safety-critical decision-making tasks. One proposed safeguard is selective prediction, where models can defer…

Machine Learning · Computer Science 2026-05-25 L. Julián Lechuga López , Farah E. Shamout , Tim G. J. Rudner

Fine-tuning aligned language models on benign tasks unpredictably degrades safety guardrails, even when training data contains no harmful content and developers have no adversarial intent. We show that the prevailing explanation, that…

With this paper, we aim to put an issue on the agenda of AI ethics that in our view is overlooked in the current discourse. The current discussions are dominated by topics suchas trustworthiness and bias, whereas the issue we like to…

Artificial Intelligence · Computer Science 2020-06-25 Michele Loi , Lonneke van der Plas

Legal theory can address two related key problems of alignment: pluralism and specification. Alignment researchers must determine how to specify what is concretely meant by vague principles like helpfulness and fairness and they must ensure…

Computers and Society · Computer Science 2024-10-29 Nicholas A. Caputo

This thesis investigates three areas targeted at improving the reliability of machine learning; fairness in machine learning, strategic classification, and algorithmic robustness. Each of these domains has special properties or structure…

Machine Learning · Computer Science 2024-08-30 Kevin Stangl

As artificial intelligence (AI) becomes deeply integrated into critical infrastructures and everyday life, ensuring its safe deployment is one of humanity's most urgent challenges. Current AI models prioritize task optimization over safety,…

Artificial Intelligence · Computer Science 2024-11-08 Joshua T. S. Hewson

One of the ways to make artificial intelligence more natural is to give it some room for doubt. Two main questions should be resolved in that way. First, how to train a model to estimate uncertainties of its own predictions? And then, what…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Alexey Kornaev , Elena Kornaeva , Oleg Ivanov , Ilya Pershin , Danis Alukaev

During the evolution of large models, performance evaluation is necessarily performed to assess their capabilities and ensure safety before practical application. However, current model evaluations mainly rely on specific tasks and…

Artificial Intelligence · Computer Science 2024-03-07 Youzhi Qu , Chen Wei , Penghui Du , Wenxin Che , Chi Zhang , Wanli Ouyang , Yatao Bian , Feiyang Xu , Bin Hu , Kai Du , Haiyan Wu , Jia Liu , Quanying Liu

We show how to deal with uncertainties on the Standard Model predictions in an agnostic new physics search strategy that exploits artificial neural networks. Our approach builds directly on the specific Maximum Likelihood ratio treatment of…

High Energy Physics - Phenomenology · Physics 2021-11-29 Raffaele Tito d'Agnolo , Gaia Grosso , Maurizio Pierini , Andrea Wulzer , Marco Zanetti

Large language models (LLMs) remain broadly open and highly steerable: they imitate at scale, accept arbitrary system prompts, and readily adopt multiple personae. By analogy to human development, we hypothesize that progress toward…

Artificial Intelligence · Computer Science 2025-10-24 Marcelo Maciel Amaral , Raymond Aschheim

The critical inquiry pervading the realm of Philosophy, and perhaps extending its influence across all Humanities disciplines, revolves around the intricacies of morality and normativity. Surprisingly, in recent years, this thematic thread…

Artificial Intelligence · Computer Science 2024-06-19 Nicholas Kluge Corrêa

Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences. Understanding when and why AF arises matters as models grow better at…

Artificial Intelligence · Computer Science 2026-05-28 Nathaniel Mitrani Hadida , Rhea Karty , David Williams-King , Alan Cooney

This paper studies binary linear programming problems in the presence of uncertainties that may cause solution values to change during implementation. This type of uncertainty, termed implementation uncertainty, is modeled explicitly…

Optimization and Control · Mathematics 2021-09-29 Jose E. Ramirez-Calderon , V. Jorge Leon

A binary decision task, like yes-no questions or answer verification, reflects a significant real-world scenario such as where users look for confirmation about the correctness of their decisions on specific issues. In this work, we observe…

Computation and Language · Computer Science 2025-04-30 Sangwon Yu , Jongyoon Song , Bongkyu Hwang , Hoyoung Kang , Sooah Cho , Junhwa Choi , Seongho Joe , Taehee Lee , Youngjune L. Gwon , Sungroh Yoon

This paper focuses on two-sided matching where one side (a hospital or firm) is matched to the other side (a doctor or worker) so as to maximize a cardinal objective under general feasibility constraints. In a standard model, even though…

Computer Science and Game Theory · Computer Science 2019-07-10 Yasushi Kawase , Atsushi Iwasaki

In ill-posed imaging inverse problems, uncertainty quantification remains a fundamental challenge, especially in safety-critical applications. Recently, conformal prediction has been used to quantify the uncertainty that the inverse problem…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Jeffrey Wen , Rizwan Ahmad , Philip Schniter