中文
相关论文

相关论文: A logical alarm for misaligned binary classifiers

200 篇论文

We study the finite satisfiability problem for the two-variable fragment of first-order logic extended with counting quantifiers (C2) and interpreted over linearly ordered structures. We show that the problem is undecidable in the case of…

计算机科学中的逻辑 · 计算机科学 2019-03-14 Witold Charatonik , Piotr Witkowski

To build robust, fair, and safe AI systems, we would like our classifiers to say ``I don't know'' when facing test examples that are difficult or fall outside of the training classes.The ubiquitous strategy to predict under uncertainty is…

机器学习 · 统计学 2024-01-22 Kamalika Chaudhuri , David Lopez-Paz

Large language models (LLMs) are increasingly proposed as agents in strategic decision environments, yet their behavior in structured geopolitical simulations remains under-researched. We evaluate six popular state-of-the-art LLMs alongside…

计算与语言 · 计算机科学 2026-03-03 Veronika Solopova , Viktoria Skorik , Maksym Tereshchenko , Alina Haidun , Ostap Vykhopen

Recent advances in AI research make it increasingly plausible that artificial agents with consequential real-world impact will soon operate beyond tightly controlled environments. Ensuring that these agents are not only safe but that they…

计算机与社会 · 计算机科学 2025-06-10 Kevin Baum

Social biases and belief-driven behaviors can significantly impact Large Language Models (LLMs) decisions on several tasks. As LLMs are increasingly used in multi-agent systems for societal simulations, their ability to model fundamental…

计算与语言 · 计算机科学 2025-10-09 Angana Borah , Marwa Houalla , Rada Mihalcea

This paper considers the estimation of binary choice models when survey responses are possibly misclassified but one of the response category can be validated. Partial validation may occur when survey questions about participation include…

计量经济学 · 经济学 2025-12-17 Augustine Denteh , Pierre E. Nguimkeu

We derive an accounting identity for predictive models that links accuracy with common fairness criteria. The identity shows that for globally calibrated models, the weighted sums of miscalibration within groups and error imbalance across…

机器学习 · 计算机科学 2026-01-29 Hadi Elzayn , Jacob Goldin

Proper scoring rules elicit truth-telling when making predictions, or otherwise revealing information. However, when multiple predictions are made of the same event, telling the truth is in general no longer optimal, as agents are motivated…

计算机科学与博弈论 · 计算机科学 2017-07-04 Amir Ban

The gold standard in human-AI collaboration is complementarity -- when combined performance exceeds both the human and algorithm alone. We investigate this challenge in binary classification settings where the goal is to maximize 0-1…

人工智能 · 计算机科学 2024-11-26 Kenny Peng , Nikhil Garg , Jon Kleinberg

In this paper, we study the accuracy of values aggregated over classes predicted by a classification algorithm. The problem is that the resulting aggregates (e.g., sums of a variable) are known to be biased. The bias can be large even for…

机器学习 · 统计学 2019-12-02 Q. A. Meertens , C. G. H. Diks , H. J. van den Herik , F W Takes

Multi-class classification methods that produce sets of probabilistic classifiers, such as ensemble learning methods, are able to model aleatoric and epistemic uncertainty. Aleatoric uncertainty is then typically quantified via the Bayes…

机器学习 · 统计学 2023-04-20 Thomas Mortier , Viktor Bengs , Eyke Hüllermeier , Stijn Luca , Willem Waegeman

The growing awareness of safety concerns in large language models (LLMs) has sparked considerable interest in the evaluation of safety. This study investigates an under-explored issue about the evaluation of LLMs, namely the substantial…

计算与语言 · 计算机科学 2024-04-02 Yixu Wang , Yan Teng , Kexin Huang , Chengqi Lyu , Songyang Zhang , Wenwei Zhang , Xingjun Ma , Yu-Gang Jiang , Yu Qiao , Yingchun Wang

Collective phenomena in systems of interacting agents have helped us understand diverse social, ecological and biological observations. The corresponding explanations are challenged by incorrect information processing. In particular, the…

物理与社会 · 物理学 2022-04-08 Johannes Falk , Edwin Eichler , Katja Windt , Marc-Thorsten Hütt

The growing adoption of large language models in legal practice brings both significant promise and serious risk. Legal professionals stand to benefit from AI that can reason over contracts, draft documents, and analyze sources at scale,…

人工智能 · 计算机科学 2026-05-15 Olivia Peiyu Wang , Leilani H. Gilpin

In this paper, we discuss different models for human logic systems and describe a game with nature. Godel`s incompleteness theorem is taken into account to construct a model of logical networks based on axioms obtained by symmetry breaking.…

物理与社会 · 物理学 2007-05-23 Fariel Shafee

Recent advances in Large Language Models (LLMs) have enabled multi-agent systems that simulate real-world interactions with near-human reasoning. While previous studies have extensively examined biases related to protected attributes such…

人工智能 · 计算机科学 2025-06-03 Min Choi , Keonwoo Kim , Sungwon Chae , Sangyeob Baek

In two-player cooperative games, agents can play together effectively when they have accurate assumptions about how their teammate will behave, but may perform poorly when these assumptions are inaccurate. In language games, failure may be…

人工智能 · 计算机科学 2024-12-18 Joseph Bills , Christopher Archibald , Diego Blaylock

Despite the impressive performance of large language models (LLMs) across various benchmarks, their ability to address ambiguously specified problems--frequent in real-world interactions--remains underexplored. To address this gap, we…

计算与语言 · 计算机科学 2025-02-10 Katarzyna Kobalczyk , Nicolas Astorga , Tennison Liu , Mihaela van der Schaar

Uncertainty estimation is a significant issue for current large language models (LLMs) that are generally poorly calibrated and over-confident, especially with reinforcement learning from human feedback (RLHF). Unlike humans, whose…

计算与语言 · 计算机科学 2024-05-13 Ruixin Yang , Dheeraj Rajagopal , Shirley Anugrah Hayati , Bin Hu , Dongyeop Kang

Structured prediction is ubiquitous in applications of machine learning such as knowledge extraction and natural language processing. Structure often can be formulated in terms of logical constraints. We consider the question of how to…

人工智能 · 计算机科学 2017-09-27 Emmanouil Antonios Platanios , Ashish Kapoor , Eric Horvitz