中文
相关论文

相关论文: A logical alarm for misaligned binary classifiers

200 篇论文

When should we delegate decisions to AI systems? While the value alignment literature has developed techniques for shaping AI values, less attention has been paid to how to determine, under uncertainty, when imperfect alignment is good…

人工智能 · 计算机科学 2025-12-23 Daniel A. Herrmann , Abinav Chari , Isabelle Qian , Sree Sharvesh , B. A. Levinstein

To collect large scale annotated data, it is inevitable to introduce label noise, i.e., incorrect class labels. To be robust against label noise, many successful methods rely on the noisy classifiers (i.e., models trained on the noisy…

计算机视觉与模式识别 · 计算机科学 2020-11-23 Songzhu Zheng , Pengxiang Wu , Aman Goswami , Mayank Goswami , Dimitris Metaxas , Chao Chen

We consider the problem of distribution-free conformal prediction and the criterion of group conditional validity. This criterion is motivated by many practical scenarios including hidden stratification and group fairness. Existing methods…

机器学习 · 计算机科学 2023-03-21 Samuel Deng , Navid Ardeshir , Daniel Hsu

Is there an equilibrium for distributed consensus when all agents except one collude to steer the decision value towards their preference? If an equilibrium exists, then an $n-1$ size coalition cannot do better by deviating from the…

分布式、并行与集群计算 · 计算机科学 2019-09-10 Yehuda Afek , Itay Harel , Amit Jacob-Fanani , Moshe Sulamy

Effective human-AI collaboration requires a system design that provides humans with meaningful ways to make sense of and critically evaluate algorithmic recommendations. In this paper, we propose a way to augment human-AI collaboration by…

机器学习 · 计算机科学 2022-05-03 Maria De-Arteaga , Alexandra Chouldechova , Artur Dubrawski

For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for language agents, arising from accidental misspecification by…

人工智能 · 计算机科学 2021-03-30 Zachary Kenton , Tom Everitt , Laura Weidinger , Iason Gabriel , Vladimir Mikulik , Geoffrey Irving

Computational argumentation offers formal frameworks for transparent, verifiable reasoning but has traditionally been limited by its reliance on domain-specific information and extensive feature engineering. In contrast, LLMs excel at…

人工智能 · 计算机科学 2026-03-18 Stylianos Loukas Vasileiou , Antonio Rago , Francesca Toni , William Yeoh

In group testing, the task is to identify defective items by testing groups of them together using as few tests as possible. We consider the setting where each item is defective with a constant probability $\alpha$, independent of all other…

离散数学 · 计算机科学 2024-11-15 Lukas Hintze , Lena Krieg , Olga Scheftelowitsch , Haodong Zhu

We make two contributions to the problem of estimating the $L_1$ calibration error of a binary classifier from a finite dataset. First, we provide an upper bound for any classifier where the calibration function has bounded variation.…

We develop a new approach to multi-label conformal prediction in which we aim to output a precise set of promising prediction candidates with a bounded number of incorrect answers. Standard conformal prediction provides the ability to adapt…

机器学习 · 计算机科学 2022-02-16 Adam Fisch , Tal Schuster , Tommi Jaakkola , Regina Barzilay

Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-variance accuracy estimates and unstable rankings across platforms. On difficult problems, an LLM may fail…

机器学习 · 计算机科学 2026-02-04 Zihan Dong , Zhixian Zhang , Yang Zhou , Can Jin , Ruijia Wu , Linjun Zhang

Tabular anomaly detection is often handled by single detectors or static ensembles, even though strong performance on tabular data typically comes from heterogeneous model families (e.g., tree ensembles, deep tabular networks, and tabular…

机器学习 · 计算机科学 2026-02-17 Pinqiao Wang , Sheng Li

We introduce a novel non-cooperative game to analyse opinion formation and resistance, incorporating principles from social psychology such as confirmation bias, resource constraints, and influence penalties. Our simulation features Large…

人工智能 · 计算机科学 2025-09-03 Amin Qasmi , Usman Naseem , Mehwish Nasim

Artificial Intelligence (AI) systems are increasingly used in high-stakes domains of our life, increasing the need to explain these decisions and to make sure that they are aligned with how we want the decision to be made. The field of…

人工智能 · 计算机科学 2023-06-28 Sofie Goethals , David Martens , Theodoros Evgeniou

As machine learning systems become more powerful they also become increasingly unpredictable and opaque. Yet, finding human-understandable explanations of how they work is essential for their safe deployment. This technical report…

Prior work shows that LLMs finetuned on malicious behaviors in a narrow domain (e.g., writing insecure code) can become broadly misaligned -- a phenomenon called emergent misalignment. We investigate whether this extends from conventional…

机器学习 · 计算机科学 2025-07-11 James Chua , Jan Betley , Mia Taylor , Owain Evans

In recent progress, mathematical verifiers have achieved success in mathematical reasoning tasks by validating the correctness of solutions generated by policy models. However, existing verifiers are trained with binary classification…

计算与语言 · 计算机科学 2024-10-21 Bofei Gao , Zefan Cai , Runxin Xu , Peiyi Wang , Ce Zheng , Runji Lin , Keming Lu , Dayiheng Liu , Chang Zhou , Wen Xiao , Junjie Hu , Tianyu Liu , Baobao Chang

As artificial agents become increasingly capable, what internal structure is *necessary* for an agent to act competently under uncertainty? Classical results show that optimal control can be *implemented* using belief states or world…

机器学习 · 计算机科学 2026-04-03 Aran Nayebi

Rule based classifiers that use the presence and absence of key sub-strings to make classification decisions have a natural mechanism for quantifying the uncertainty of their precision. For a binary classifier, the key insight is to treat…

机器学习 · 计算机科学 2020-05-20 James Nutaro , Ozgur Ozmen

A principal and an agent can launch a project under unanimous consent. Their individual payoffs from the project depend on an underlying state, and the agent privately knows his own preference. The principal can conduct a test to learn…

理论经济学 · 经济学 2026-02-06 Yingkai Li , Boli Xu