中文
相关论文

相关论文: ProPublica's COMPAS Data Revisited

200 篇论文

We propose a novel approach for inferring the individualized causal effects of a treatment (intervention) from observational data. Our approach conceptualizes causal inference as a multitask learning problem; we model a subject's potential…

机器学习 · 计算机科学 2017-06-20 Ahmed M. Alaa , Michael Weisz , Mihaela van der Schaar

We restrict the propagation of misinformation in a social-media-like environment while preserving the spread of correct information. We model the environment as a random network of users in which each news item propagates in the network in…

社会与信息网络 · 计算机科学 2022-11-10 Yigit E. Bayiz , Ufuk Topcu

Selection problems with costly information, dating back to Weitzman's Pandora's Box problem, have received much attention recently. We study the general model of Costly Information Combinatorial Selection (CICS) that was recently introduced…

数据结构与算法 · 计算机科学 2025-12-09 Shuchi Chawla , Dimitris Christou , Trung Dang

Biases in existing datasets used to train algorithmic decision rules can raise ethical and economic concerns due to the resulting disparate treatment of different groups. We propose an algorithm for sequentially debiasing such datasets…

机器学习 · 计算机科学 2023-01-11 Yifan Yang , Yang Liu , Parinaz Naghizadeh

Selective inference (post-selection inference) is a methodology that has attracted much attention in recent years in the fields of statistics and machine learning. Naive inference based on data that are also used for model selection tends…

统计方法学 · 统计学 2021-11-25 Yoshiyuki Ninomiya , Yuta Umezu , Ichiro Takeuchi

Missing data are ubiquitous in the era of big data and, if inadequately handled, are known to lead to biased findings and have deleterious impact on data-driven decision makings. To mitigate its impact, many missing value imputation methods…

机器学习 · 计算机科学 2021-10-26 Yiliang Zhang , Qi Long

Unmeasured confounding and selection bias are often of concern in observational studies and may invalidate a causal analysis if not appropriately accounted for. Under outcome-dependent sampling, a latent factor that has causal effects on…

统计方法学 · 统计学 2022-08-03 Kendrick Qijun Li , Xu Shi , Wang Miao , Eric Tchetgen Tchetgen

Algorithmic risk assessments are increasingly used to help humans make decisions in high-stakes settings, such as medicine, criminal justice and education. In each of these cases, the purpose of the risk assessment tool is to inform…

机器学习 · 统计学 2020-01-13 Amanda Coston , Alan Mishler , Edward H. Kennedy , Alexandra Chouldechova

The presence of confounding by high-dimensional variables complicates estimation of the average effect of a point treatment. On the one hand, it necessitates the use of variable selection strategies or more general data-adaptive…

统计方法学 · 统计学 2017-08-15 Vahe Avagyan , Stijn Vansteelandt

Data-driven learning algorithms are employed in many online applications, in which data become available over time, like network monitoring, stock price prediction, job applications, etc. The underlying data distribution might evolve over…

机器学习 · 计算机科学 2021-08-16 Vasileios Iosifidis , Wenbin Zhang , Eirini Ntoutsi

Face recognition algorithms, when used in the real world, can be very useful, but they can also be dangerous when biased toward certain demographics. So, it is essential to understand how these algorithms are trained and what factors affect…

计算机视觉与模式识别 · 计算机科学 2023-02-14 Manideep Kolla , Aravinth Savadamuthu

As predictive algorithms grow in popularity, using the same dataset to both train and test a new model has become routine across research, policy, and industry. Sample-splitting attains valid inference on model properties by using separate…

计量经济学 · 经济学 2025-11-27 Bruno Fava

Progressive censoring scheme has received considerable attention in recent years. In this paper we introduce a new type-II progressive censoring scheme for two samples. It is observed that the proposed censoring scheme is analytically more…

统计方法学 · 统计学 2016-09-20 Shuvashree Mondal , Debasis Kundu

Recently, Thas et al. (2012) introduced a new statistical model for the probability index. This index is defined as $P(Y \leq Y^*|X, X^*)$ where Y and Y* are independent random response variables associated with covariates X and X* [...]…

统计计算 · 统计学 2018-08-20 Han Bossier , Gustavo Amorim , Jan De Neve , Olivier Thas

Evaluating fairness can be challenging in practice because the sensitive attributes of data are often inaccessible due to privacy constraints. The go-to approach that the industry frequently adopts is using off-the-shelf proxy models to…

机器学习 · 计算机科学 2023-02-01 Zhaowei Zhu , Yuanshun Yao , Jiankai Sun , Hang Li , Yang Liu

Propensity scores are commonly used to reduce the confounding bias in non-randomized observational studies for estimating the average treatment effect. An important assumption underlying this approach is that all confounders that are…

统计方法学 · 统计学 2022-08-02 Youfei Yu , Jiacong Du , Min Zhang , Zhenke Wu , Andrew M. Ryan , Bhramar Mukherjee

With the rapid growth in language processing applications, fairness has emerged as an important consideration in data-driven solutions. Although various fairness definitions have been explored in the recent literature, there is lack of…

机器学习 · 计算机科学 2022-03-17 Satyapriya Krishna , Rahul Gupta , Apurv Verma , Jwala Dhamala , Yada Pruksachatkun , Kai-Wei Chang

Scholars have focused on algorithms used during sentencing, bail, and parole, but little work explores what we call carceral algorithms that are used during incarceration. This paper is focused on the Pennsylvania Additive Classification…

计算机与社会 · 计算机科学 2021-12-07 Vanessa Massaro , Swarup Dhar , Darakhshan Mir , Nathan C. Ryan

We study the problem of reconstructing tabular data from aggregate statistics, in which the attacker aims to identify interesting claims about the sensitive data that can be verified with 100% certainty given the aggregates. Successful…

机器学习 · 统计学 2025-06-12 Terrance Liu , Eileen Xiao , Adam Smith , Pratiksha Thaker , Zhiwei Steven Wu

Its crux lies in the optimization of a tradeoff between accuracy and fairness of resultant models on the selected feature subset. The technical challenge of our setting is twofold: 1) streaming feature inputs, such that an informative…

机器学习 · 计算机科学 2024-08-26 Leizhen Zhang , Lusi Li , Di Wu , Sheng Chen , Yi He