English
Related papers

Related papers: Lurking Inferential Monsters? Quantifying bias in …

200 papers

This paper shows that two commonly used evaluation metrics for generative models, the Fr\'echet Inception Distance (FID) and the Inception Score (IS), are biased -- the expected value of the score computed for a finite sample set is not the…

Computer Vision and Pattern Recognition · Computer Science 2020-06-17 Min Jin Chong , David Forsyth

In designed experiments and surveys, known laws or design feat ures provide checks on the most relevant aspects of a model and identify the target parameters. In contrast, in most observational studies in the health and social sciences, the…

Methodology · Statistics 2010-01-18 Sander Greenland

A commonly observed pattern in machine learning models is an underprediction of the target feature, with the model's predicted target rate for members of a given category typically being lower than the actual target rate for members of that…

Machine Learning · Computer Science 2023-07-06 Owen O'Neill , Fintan Costello

Addressing selection bias in latent variable causal discovery is important yet underexplored, largely due to a lack of suitable statistical tools: While various tools beyond basic conditional independencies have been developed to handle…

Machine Learning · Computer Science 2025-12-15 Haoyue Dai , Yiwen Qiu , Ignavier Ng , Xinshuai Dong , Peter Spirtes , Kun Zhang

A key element in transfer learning is representation learning; if representations can be developed that expose the relevant factors underlying the data, then new tasks and domains can be learned readily based on mappings of these salient…

Machine Learning · Computer Science 2014-12-18 Yujia Li , Kevin Swersky , Richard Zemel

This paper addresses the critical issue of sample selection bias in cross-country comparisons based on international assessments such as the Programme for International Student Assessment (PISA). Although PISA is widely used to benchmark…

Econometrics · Economics 2025-10-08 Onil Boussim

With the resurgence of interest in neural networks, representation learning has re-emerged as a central focus in artificial intelligence. Representation learning refers to the discovery of useful encodings of data that make domain-relevant…

Machine Learning · Computer Science 2016-12-19 Karl Ridgeway

Sense embedding learning methods learn different embeddings for the different senses of an ambiguous word. One sense of an ambiguous word might be socially biased while its other senses remain unbiased. In comparison to the numerous prior…

Computation and Language · Computer Science 2022-03-17 Yi Zhou , Masahiro Kaneko , Danushka Bollegala

We consider the problem of sequential evaluation, in which an evaluator observes candidates in a sequence and assigns scores to these candidates in an online, irrevocable fashion. Motivated by the psychology literature that has studied…

Machine Learning · Statistics 2023-11-20 Jingyan Wang , Ashwin Pananjady

We consider large-scale studies in which it is of interest to test a very large number of hypotheses, and then to estimate the effect sizes corresponding to the rejected hypotheses. For instance, this setting arises in the analysis of gene…

Methodology · Statistics 2015-03-31 Kean Ming Tan , Noah Simon , Daniela Witten

Consider a finite sample from an unknown distribution over a countable alphabet. Unobserved events are alphabet symbols which do not appear in the sample. Estimating the probabilities of unobserved events is a basic problem in statistics…

Statistics Theory · Mathematics 2022-11-08 Amichai Painsky

The missing mass refers to the proportion of data points in an unknown population of classifier inputs that belong to classes not present in the classifier's training data, which is assumed to be a random sample from that unknown…

Machine Learning · Computer Science 2025-03-11 Seongmin Lee , Marcel Böhme

Are certain cognitive biases mathematically inevitable consequences of sequential information processing? We prove that primacy effects, anchoring, and order-dependence are architecturally necessary in autoregressive language models due to…

Artificial Intelligence · Computer Science 2026-05-12 Jikun Wu , Dongxin Guo , Siu-Ming Yiu

Bias in Large Language Models (LLMs) significantly undermines their reliability and fairness. We focus on a common form of bias: when two reference concepts in the model's concept space, such as sentiment polarities (e.g., "positive" and…

Computation and Language · Computer Science 2025-05-22 Lang Gao , Kaiyang Wan , Wei Liu , Chenxi Wang , Zirui Song , Zixiang Xu , Yanbo Wang , Veselin Stoyanov , Xiuying Chen

We develop a novel method for personalized off-policy learning in scenarios with unobserved confounding. Thereby, we address a key limitation of standard policy learning: standard policy learning assumes unconfoundedness, meaning that no…

Machine Learning · Computer Science 2026-02-18 Konstantin Hess , Dennis Frauen , Valentyn Melnychuk , Stefan Feuerriegel

Latent variable models are popularly used to measure latent factors (e.g., abilities and personalities) from large-scale assessment data. Beyond understanding these latent factors, the covariate effect on responses controlling for latent…

Methodology · Statistics 2026-01-12 Jing Ouyang , Chengyu Cui , Kean Ming Tan , Gongjun Xu

Identifying and disentangling sources of predictive uncertainty is essential for trustworthy supervised learning. We argue that widely used second-order methods that disentangle aleatoric and epistemic uncertainty are fundamentally…

Machine Learning · Computer Science 2026-02-09 Sebastián Jiménez , Mira Jürgens , Willem Waegeman

In settings where interference between units is possible, we define the prevalence of indirect effects to be the number of units who are affected by the treatment of others. This quantity does not fully identify an indirect effect, but may…

Methodology · Statistics 2024-01-18 David Choi

We consider the problem of learning fair policies for multi-stage selection problems from observational data. This problem arises in several high-stakes domains such as company hiring, loan approval, or bail decisions where outcomes (e.g.,…

Machine Learning · Computer Science 2023-12-21 Zhuangzhuang Jia , Grani A. Hanasusanto , Phebe Vayanos , Weijun Xie

A well-known problem when learning from user clicks are inherent biases prevalent in the data, such as position or trust bias. Click models are a common method for extracting information from user clicks, such as document relevance in web…

Information Retrieval · Computer Science 2024-12-17 Romain Deffayet , Philipp Hager , Jean-Michel Renders , Maarten de Rijke