Related papers: Statistical discrimination and statistical informa…
In observational studies of discrimination, the most common statistical approaches consider either the rate at which decisions are made (benchmark tests) or the success rate of those decisions (outcome tests). Both tests, however, have…
We give an overview of the role of information theory in statistics, and particularly in biostatistics. We recall the basic quantities in information theory; entropy, cross-entropy, conditional entropy, mutual information and…
This paper studies incomplete-information games in which an information provider, an oracle, publicly discloses information to the players. One oracle is said to dominate another if, in every game, it can replicate the equilibrium outcomes…
We study the problem of sequential prediction of categorical data and discuss a generalisation of Blackwell's algorithm on 0-1 data. The arguments are based on Blackwell's approachability results given in Blackwell, D. (1956): An Analog of…
We offer a search-theoretic model of statistical discrimination, in which firms treat identical groups unequally based on their occupational choices. The model admits symmetric equilibria in which the group characteristic is ignored, but…
Algorithmic discrimination is an important aspect when data is used for predictive purposes. This paper analyzes the relationships between discrimination and classification, data set partitioning, and decision models, as well as…
This paper provides a unified approach to characterize the set of all feasible signals subject to privacy constraints. The Blackwell frontier of feasible signals can be decomposed into minimum informative signals achieving the Blackwell…
We show that the ability to consider counterfactual situations is a necessary assumption of Bell's theorem, and that, to allow Bell inequality violations while maintaining all other assumptions, we just require certain measurement choices…
An analyst is tasked with producing a statistical study. The analyst is not monitored and is able to manipulate the study. He can receive payments contingent on his report and trusted data collected from an independent source, modeled as a…
It is well known that the independence of the sample mean and the sample variance characterizes the normal distribution. By using Anosov's theorem, we further investigate the analogous characteristic properties in terms of the sample mean…
Data plays a central role in advancements in modern artificial intelligence, with high-quality data emerging as a key driver of model performance. This has prompted the development of principled and effective data curation methods in recent…
Information disclosure can compromise privacy when revealed information is correlated with private information. We consider the notion of inferential privacy, which measures privacy leakage by bounding the inferential power a Bayesian…
When a machine-learning algorithm makes biased decisions, it can be helpful to understand the sources of disparity to explain why the bias exists. Towards this, we examine the problem of quantifying the contribution of each individual…
I study statistical discrimination driven by verifiable beliefs, such as those generated by machine learning, rather than by humans. When beliefs are verifiable, interventions against statistical discrimination can move beyond simple,…
This study extends Blackwell's (1953) comparison of information to a sequential social learning model, where agents make decisions sequentially based on both private signals and the observed actions of others. In this context, we introduce…
The inference of causal relationships using observational data from partially observed multivariate systems with hidden variables is a fundamental question in many scientific domains. Methods extracting causal information from conditional…
Controversy about the significance of underdetermination of theories persists in the philosophy and conduct of science. The issue has practical import when research is used to inform decision making, because scientific uncertainty yields…
Causal analysis may be affected by selection bias, which is defined as the systematic exclusion of data from a certain subpopulation. Previous work in this area focused on the derivation of identifiability conditions. We propose instead a…
Renyi's "thinning" operation on a discrete random variable is a natural discrete analog of the scaling operation for continuous random variables. The properties of thinning are investigated in an information-theoretic context, especially in…
Discrimination discovery from data is an important task aiming at identifying patterns of illegal and unethical discriminatory activities against protected-by-law groups, e.g., ethnic minorities. While any legally-valid proof of…