中文
相关论文

相关论文: Lurking Inferential Monsters? Quantifying bias in …

200 篇论文

In this paper we study covariance estimation with missing data. We consider missing data mechanisms that can be independent of the data, or have a time varying dependency. Additionally, observed variables may have arbitrary (non uniform)…

统计理论 · 数学 2021-06-17 Eduardo Pavez , Antonio Ortega

We present new results on average causal effects in settings with unmeasured exposure-outcome confounding. Our results are motivated by a class of estimands, e.g., frequently of interest in medicine and public health, that are currently not…

统计方法学 · 统计学 2023-12-25 Lan Wen , Aaron L. Sarvet , Mats J. Stensrud

Accurate estimates of examination bias are crucial for unbiased learning-to-rank from implicit feedback in search engines and recommender systems, since they enable the use of Inverse Propensity Score (IPS) weighting techniques to address…

信息检索 · 计算机科学 2019-05-27 Zhichong Fang , Aman Agarwal , Thorsten Joachims

Negative binomial regression is commonly employed to analyze overdispersed count data. With small to moderate sample sizes, the maximum likelihood estimator of the dispersion parameter may be subject to a significant bias, that in turn…

统计方法学 · 统计学 2020-11-06 Euloge Clovis Kenne Pagui , Alessandra Salvan , Nicola Sartori

One of the difficulties of artificial intelligence is to ensure that model decisions are fair and free of bias. In research, datasets, metrics, techniques, and tools are applied to detect and mitigate algorithmic unfairness and bias. This…

Scoring models support decision-making in financial institutions. Their estimation and evaluation are based on the data of previously accepted applicants with known repayment behavior. This creates sampling bias: the available labeled data…

As language technologies become widespread, it is important to understand how changes in language affect reader perceptions and behaviors. These relationships may be formalized as the isolated causal effect of some focal language-encoded…

计算与语言 · 计算机科学 2025-06-06 Victoria Lin , Louis-Philippe Morency , Eli Ben-Michael

We argue that the selective inclusion of data points based on latent objectives is common in practical situations, such as music sequences. Since this selection process often distorts statistical analysis, previous work primarily views it…

机器学习 · 计算机科学 2024-07-02 Yujia Zheng , Zeyu Tang , Yiwen Qiu , Bernhard Schölkopf , Kun Zhang

Selective classification is a powerful tool for automated decision-making in high-risk scenarios, allowing classifiers to act only when confident and abstain when uncertainty is high. Given a target accuracy, our goal is to minimize…

统计理论 · 数学 2025-10-28 Mohamed Ndaoud , Peter Radchenko , Bradley Rava

Implicit bias is the unconscious attribution of particular qualities (or lack thereof) to a member from a particular social group (e.g., defined by gender or race). Studies on implicit bias have shown that these unconscious stereotypes can…

计算机与社会 · 计算机科学 2020-01-27 L. Elisa Celis , Anay Mehrotra , Nisheeth K. Vishnoi

The lack of non-parametric statistical tests for confounding bias significantly hampers the development of robust, valid and generalizable predictive models in many fields of research. Here I propose the partial and full confounder tests,…

机器学习 · 计算机科学 2025-05-30 Tamas Spisak

Discrimination can occur when the underlying unbiased labels are overwritten by an agent with potential bias, resulting in biased datasets that unfairly harm specific groups and cause classifiers to inherit these biases. In this paper, we…

机器学习 · 计算机科学 2023-12-27 Yixuan Zhang , Boyu Li , Zenan Ling , Feng Zhou

We establish concentration rates for estimation of treatment effects in experiments that incorporate prior sources of information -- such as past pilots, related studies, or expert assessments -- whose external validity is uncertain. Each…

计量经济学 · 经济学 2026-03-24 Frederico Finan , Demian Pouzo

This paper studies the semi-parametric identification and estimation of a rational inattention model with Bayesian persuasion. The identification requires the observation of a cross-section of market-level outcomes. The empirical content of…

计量经济学 · 经济学 2020-09-18 Moyu Liao

Automated Essay Scoring (AES) has been quite popular and is being widely used. However, lack of appropriate methodology for rating nonnative English speakers' essays has meant a lopsided advancement in this field. In this paper, we report…

计算与语言 · 计算机科学 2018-02-02 Amber Nigam

Increasingly large parameter spaces, used to more accurately model precision observables in physics, can paradoxically lead to large deviations in the inferred parameters of interest -- a bias known as volume projection effects -- when…

宇宙学与河外天体物理 · 物理学 2025-07-29 Alexander Reeves , Pierre Zhang , Henry Zheng

A primary difficulty with unsupervised discovery of structure in large data sets is a lack of quantitative evaluation criteria. In this work, we propose and investigate several metrics for evaluating and comparing generative models of…

机器学习 · 计算机科学 2020-07-27 Daniel Jiwoong Im , Iljung Kwak , Kristin Branson

Model selection is a necessary step in unsupervised machine learning. Despite numerous criteria and metrics, model selection remains subjective. A high degree of subjectivity may lead to questions about repeatability and reproducibility of…

机器学习 · 计算机科学 2024-01-08 Wanyi Chen , Mary L. Cummings

We investigate the problem of reliably assessing group fairness when labeled examples are few but unlabeled examples are plentiful. We propose a general Bayesian framework that can augment labeled data with unlabeled data to produce more…

机器学习 · 统计学 2020-10-21 Disi Ji , Padhraic Smyth , Mark Steyvers

Selective inference (post-selection inference) is a methodology that has attracted much attention in recent years in the fields of statistics and machine learning. Naive inference based on data that are also used for model selection tends…

统计方法学 · 统计学 2021-11-25 Yoshiyuki Ninomiya , Yuta Umezu , Ichiro Takeuchi