English
Related papers

Related papers: Data selection and confounding in the court case o…

200 papers

Non-discrimination is a recognized objective in algorithmic decision making. In this paper, we introduce a novel probabilistic formulation of data pre-processing for reducing discrimination. We propose a convex optimization for learning a…

Machine Learning · Statistics 2017-04-12 Flavio P. Calmon , Dennis Wei , Karthikeyan Natesan Ramamurthy , Kush R. Varshney

Over the years, there has been growing interest in using Machine Learning techniques for biomedical data processing. When tackling these tasks, one needs to bear in mind that biomedical data depends on a variety of characteristics, such as…

Machine Learning · Computer Science 2020-02-05 Elisa Ferrari , Alessandra Retico , Davide Bacciu

In many stochastic service systems, decision-makers find themselves making a sequence of decisions, with the number of decisions being unpredictable. To enhance these decisions, it is crucial to uncover the causal impact these decisions…

Methodology · Statistics 2023-07-18 Juan C. David Gomez , Amy L. Cochran , Gabriel Zayas-Caban

We revisit logistic regression and its nonlinear extensions, including multilayer feedforward neural networks, by showing that these classifiers can be viewed as converting input or higher-level features into Dempster-Shafer mass functions…

Machine Learning · Computer Science 2019-12-13 Thierry Denoeux

$\textbf{Objective}$ Develop an automatic diagnostic system which only uses textual admission information from Electronic Health Records (EHRs) and assist clinicians with a timely and statistically proved decision tool. The hope is that the…

Computation and Language · Computer Science 2017-12-08 Christy Li , Dimitris Konomis , Graham Neubig , Pengtao Xie , Carol Cheng , Eric Xing

Statistical fairness stipulates equivalent outcomes for every protected group, whereas causal fairness prescribes that a model makes the same prediction for an individual regardless of their protected characteristics. Counterfactual data…

Computation and Language · Computer Science 2024-04-02 Hannah Chen , Yangfeng Ji , David Evans

Insufficiently precise diagnosis of clinical disease is likely responsible for many treatment failures, even for common conditions and treatments. With a large enough dataset, it may be possible to use unsupervised machine learning to…

This study introduces a new approach to addressing positive and unlabeled (PU) data through the double exponential tilting model (DETM). Traditional methods often fall short because they only apply to selected completely at random (SCAR) PU…

Methodology · Statistics 2025-02-25 Siyan Liu , Chi-Kuang Yeh , Xin Zhang , Qinglong Tian , Pengfei Li

Owing to the cross-pollination between causal discovery and deep learning, non-statistical data (e.g., images, text, etc.) encounters significant conflicts in terms of properties and methods with traditional causal data. To unify these data…

Machine Learning · Computer Science 2023-08-14 Hang Chen , Xinyu Yang , Qing Yang

This paper develops a method to conduct causal inference in the presence of unobserved confounders by leveraging networks with homophily, a frequently observed tendency to form edges with similar nodes. I introduce a concept of asymptotic…

Econometrics · Economics 2025-11-04 Vincent Starck

Over the last decade, proliferation of various online platforms and their increasing adoption by billions of users have heightened the privacy risk of a user enormously. In fact, security researchers have shown that sparse microdata…

Machine Learning · Computer Science 2017-02-07 Baichuan Zhang , Noman Mohammed , Vachik Dave , Mohammad Al Hasan

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

Methodology · Statistics 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu

The basic statistical methods of data representation did not change since their emergence. Their simplicity was dictated by the intricacies of computations in the before computers epoch. It turns out that such approach is not uniquely…

Mathematical Software · Computer Science 2007-05-23 Yefim Bakman

We consider a problem of data integration. Consider determining which genes affect a disease. The genes, which we call predictor objects, can be measured in different experiments on the same individual. We address the question of finding…

Machine Learning · Statistics 2016-10-04 Xin Gao , Raymond J. Carroll

Various methods have recently been proposed to estimate causal effects with confidence intervals that are uniformly valid over a set of data generating processes when high-dimensional nuisance models are estimated by post-model-selection or…

Methodology · Statistics 2025-10-07 Niloofar Moosavi , Tetiana Gorbach , Xavier de Luna

Understanding how racial information impacts human decision making in online systems is critical in today's world. Prior work revealed that race information of criminal defendants, when presented as a text field, had no significant impact…

Human-Computer Interaction · Computer Science 2020-02-05 Keri Mallari , Kori Inkpen , Paul Johns , Sarah Tan , Divya Ramesh , Ece Kamar

In the causal adjustment setting, variable selection techniques based on one of either the outcome or treatment allocation model can result in the omission of confounders, which leads to bias, or the inclusion of spurious variables, which…

Methodology · Statistics 2015-11-30 Ashkan Ertefaie , Masoud Asgharian , David Stephens

This paper identifies the probability of causation when there is sample selection. We show that the probability of causation is partially identified for individuals who are always observed regardless of treatment status and derive sharp…

Econometrics · Economics 2024-07-08 Vitor Possebom , Flavio Riva

The focus of this paper is on the evaluation of sixteen labeling methods for hierarchical document clusters over five datasets. All of the methods are independent from clustering algorithms, applied subsequently to the dendrogram…

Information Retrieval · Computer Science 2018-05-28 Maria Fernanda Moura , Fabiano Fernandes dos Santos , Solange Oliveira Rezende

Objective: Researchers often use model-based multiple imputation to handle missing at random data to minimize bias while making the best use of all available data. However, there are sometimes constraints within the data that make…

Methodology · Statistics 2020-11-03 Chinchin Wang , Tyrel Stokes , Russell Steele , Niels Wedderkopp , Ian Shrier