English
Related papers

Related papers: Decomposition of Differences in Distribution under…

200 papers

Community detection methods attempt to divide a network into groups of nodes that share similar properties, thus revealing its large-scale structure. A major challenge when employing such methods is that they are often degenerate, typically…

Physics and Society · Physics 2021-04-23 Tiago P. Peixoto

Estimating counterfactual distributions under interventions is central to treatment risk assessment and counterfactual generation tasks. Existing approaches model the counterfactual distribution as a standalone generative target, without…

Machine Learning · Statistics 2026-05-11 Hugh Dance , Johnny Xi , Peter Orbanz , Benjamin Bloem-Reddy

Estimating individual-level treatment effect from observational data is a fundamental problem in causal inference and has attracted increasing attention in the fields of education, healthcare, and public policy.In this work, we concentrate…

Machine Learning · Computer Science 2025-07-10 Hui Meng , Keping Yang , Xuyu Peng , Bo Zheng

Strong empirical evidence from laboratory experiments, and more recently from population surveys, shows that individuals, when evaluating their situations, pay attention to whether they experience gains or losses, with losses weighing more…

Theoretical Economics · Economics 2025-10-17 Martyna Kobus , Radosław Kurek , Thomas Parker

When a model's performance differs across socially or culturally relevant groups--like race, gender, or the intersections of many such groups--it is often called "biased." While much of the work in algorithmic fairness over the last several…

Methodology · Statistics 2022-07-01 Kristian Lum , Yunfeng Zhang , Amanda Bower

We examine the impact of annual hours worked on annual earnings by decomposing changes in the real annual earnings distribution into composition, structural and hours effects. We do so via a nonseparable simultaneous model of hours, wages…

Econometrics · Economics 2021-11-19 Iván Fernández-Val , Franco Peracchi , Aico van Vuuren , Francis Vella

Difference-in-differences (diff-in-diff) is a study design that compares outcomes of two groups (treated and comparison) at two time points (pre- and post-treatment) and is widely used in evaluating new policy implementations. For instance,…

Applications · Statistics 2019-11-28 Bret Zeldow , Laura A. Hatfield

We consider a situation where the distribution of a random variable is being estimated by the empirical distribution of noisy measurements of that variable. This is common practice in, for example, teacher value-added models and other…

Econometrics · Economics 2021-12-08 Koen Jochmans , Martin Weidner

Statistical NLP systems are frequently evaluated and compared on the basis of their performances on a single split of training and test data. Results obtained using a single split are, however, subject to sampling noise. In this paper we…

Computation and Language · Computer Science 2007-05-23 Yuval Krymolowski

Selective regression allows abstention from prediction if the confidence to make an accurate prediction is not sufficient. In general, by allowing a reject option, one expects the performance of a regression model to increase at the cost of…

Machine Learning · Computer Science 2022-07-18 Abhin Shah , Yuheng Bu , Joshua Ka-Wing Lee , Subhro Das , Rameswar Panda , Prasanna Sattigeri , Gregory W. Wornell

Cooperation on social networks is crucial for understanding human survival and development. Although network structure has been found to significantly influence cooperation, human experiments have observed different cooperation phenomena…

Physics and Society · Physics 2025-08-25 Zhihao Hou , Zhikun She , Quanyi Liang , Qi Su , Daqing Li

We propose a novel framework for conducting causal inference based on counterfactual densities. While the current paradigm of causal inference is mostly focused on estimating average treatment effects (ATEs), which restricts the analysis to…

Econometrics · Economics 2026-04-27 Georg Keilbar , Sonja Greven

A common assumption in causal modeling posits that the data is generated by a set of independent mechanisms, and algorithms should aim to recover this structure. Standard unsupervised learning, however, is often concerned with training a…

Machine Learning · Computer Science 2019-03-05 Francesco Locatello , Damien Vincent , Ilya Tolstikhin , Gunnar Rätsch , Sylvain Gelly , Bernhard Schölkopf

Unequal representation of demographic groups in training data poses challenges to model generalisation across populations. Standard practice assumes that balancing subgroup representation optimises performance. However, recent empirical…

Machine Learning · Computer Science 2025-12-11 Anissa Alloula , Charles Jones , Zuzanna Wakefield-Skorniewska , Francesco Quinzan , Bartłomiej Papież

Prediction models can perform poorly when deployed to target distributions different from the training distribution. To understand these operational failure modes, we develop a method, called DIstribution Shift DEcomposition (DISDE), to…

Machine Learning · Statistics 2023-07-12 Tiffany Tianhui Cai , Hongseok Namkoong , Steve Yadlowsky

The performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-related challenges. A common issue arises when curating training data or deploying models: two…

Machine Learning · Computer Science 2025-09-24 Varun Babbar , Zhicheng Guo , Cynthia Rudin

In this paper, we extend the Riesz representation framework to causal inference under sample selection, where both treatment assignment and outcome observability are non-random. Formulating the problem in terms of a Riesz representer…

Understanding the statistical dynamics of growth and inequality is a fundamental challenge to ecology and society. Recent analyses of wealth and income dynamics in contemporary societies show that economic inequality is very dynamic and…

Physics and Society · Physics 2022-10-19 Jordan T. Kemp , Luis M. A. Bettencourt

Modeling human behavioral data is challenging due to its scale, sparseness (few observations per individual), heterogeneity (differently behaving individuals), and class imbalance (few observations of the outcome of interest). An additional…

Computers and Society · Computer Science 2018-10-24 Peter G Fennell , Zhiya Zuo , Kristina Lerman

Predictive modeling uncovers knowledge and insights regarding a hypothesized data generating mechanism (DGM). Results from different studies on a complex DGM, derived from different data sets, and using complicated models and algorithms,…

Methodology · Statistics 2022-11-21 Anna L. Smith , Tian Zheng , Andrew Gelman
‹ Prev 1 8 9 10 Next ›