English
Related papers

Related papers: How do dataset characteristics affect the performa…

200 papers

This study aims to understand how statistical biases affect the model's ability to generalize to in-distribution and out-of-distribution data on algorithmic tasks. Prior research indicates that transformers may inadvertently learn to rely…

Machine Learning · Computer Science 2024-09-11 John Mitros

Random-effects meta-analyses of observational studies can produce biased estimates if the synthesized studies are subject to unmeasured confounding. We propose sensitivity analyses quantifying the extent to which unmeasured confounding of…

Methodology · Statistics 2017-10-10 Maya B. Mathur , Tyler J. VanderWeele

Data imbalance is common in production data, where controlled production settings require data to fall within a narrow range of variation and data are collected with quality assessment in mind, rather than data analytic insights. This…

Machine Learning · Statistics 2021-12-17 Rune D. Kjærsgaard , Manja G. Grønberg , Line K. H. Clemmensen

A commonly observed pattern in machine learning models is an underprediction of the target feature, with the model's predicted target rate for members of a given category typically being lower than the actual target rate for members of that…

Machine Learning · Computer Science 2023-07-06 Owen O'Neill , Fintan Costello

Semisupervised methods inevitably invoke some assumption that links the marginal distribution of the features to the regression function of the label. Most commonly, the cluster or manifold assumptions are used which imply that the…

Statistics Theory · Mathematics 2011-12-02 Martin Azizyan , Aarti Singh , Larry Wasserman

Causal inference is crucial for understanding the true impact of interventions, policies, or actions, enabling informed decision-making and providing insights into the underlying mechanisms that shape our world. In this paper, we establish…

Methodology · Statistics 2024-03-26 Jingyue Huang , Changbao Wu , Leilei Zeng

Random-effects models are frequently used to synthesise information from different studies in meta-analysis. While likelihood-based inference is attractive both in terms of limiting properties and of implementation, its application in…

Methodology · Statistics 2018-02-16 Ioannis Kosmidis , Annamaria Guolo , Cristiano Varin

Predictive algorithms inform consequential decisions in settings with selective labels: outcomes are observed only for units selected by past decision makers. This creates an identification problem under unobserved confounding -- when…

Econometrics · Economics 2025-11-07 Ashesh Rambachan , Amanda Coston , Edward Kennedy

Large-sample data became prevalent as data acquisition became cheaper and easier. While a large sample size has theoretical advantages for many statistical methods, it presents computational challenges. Sketching, or compression, is a…

Machine Learning · Statistics 2020-05-11 Alexander F. Lapanowski , Irina Gaynanova

A key to causal inference with observational data is achieving balance in predictive features associated with each treatment type. Recent literature has explored representation learning to achieve this goal. In this work, we discuss the…

Machine Learning · Statistics 2021-02-25 Serge Assaad , Shuxi Zeng , Chenyang Tao , Shounak Datta , Nikhil Mehta , Ricardo Henao , Fan Li , Lawrence Carin

We consider the problem of selecting confounders for adjustment from a potentially large set of covariates, when estimating a causal effect. Recently, the high-dimensional Propensity Score (hdPS) method was developed for this task; hdPS…

Methodology · Statistics 2021-12-17 Asad Haris , Robert Platt

Measurement error arises through a variety of mechanisms. A rich literature exists on the bias introduced by covariate measurement error and on methods of analysis to address this bias. By comparison, less attention has been given to errors…

Methodology · Statistics 2018-11-27 Pamela Shaw , Jiwei He , Bryan Shepherd

Adaptive experiment designs can dramatically improve statistical efficiency in randomized trials, but they also complicate statistical inference. For example, it is now well known that the sample mean is biased in adaptive trials.…

Machine Learning · Statistics 2021-02-16 Vitor Hadad , David A. Hirshberg , Ruohan Zhan , Stefan Wager , Susan Athey

Comparing the internal representations of neural networks is a central goal in both neuroscience and machine learning. Standard alignment metrics operate on raw neural activations, implicitly assuming that similar representations produce…

Machine Learning · Computer Science 2026-04-02 Sunny Liu , Habon Issa , André Longon , Liv Gorton , Meenakshi Khosla , David Klindt

Distribution shifts are common in real-world datasets and can affect the performance and reliability of deep learning models. In this paper, we study two types of distribution shifts: diversity shifts, which occur when test samples exhibit…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Alceu Bissoto , Catarina Barata , Eduardo Valle , Sandra Avila

The high dimensional nature of genomics data complicates feature selection, in particular in low sample size studies - not uncommon in clinical prediction settings. It is widely recognized that complementary data on the features, `co-data',…

Methodology · Statistics 2024-05-09 Mark A. van de Wiel , Wessel N. van Wieringen

Propensity score methods were proposed by Rosenbaum and Rubin [Biometrika 70 (1983) 41--55] as central tools to help assess the causal effects of interventions. Since their introduction more than two decades ago, they have found wide…

Statistics Theory · Mathematics 2007-06-13 Donald B. Rubin , Richard P. Waterman

Diffusion-based generative models demonstrate state-of-the-art performance across various image synthesis tasks, yet their tendency to replicate and amplify dataset biases remains poorly understood. Although previous research has viewed…

Machine Learning · Computer Science 2025-12-24 Nathan Roos , Ekaterina Iakovleva , Ani Gjergji , Vito Paolo Pastore , Enzo Tartaglione

This paper studies the problem of statistical inference for genetic relatedness between binary traits based on individual-level genome-wide association data. Specifically, under the high-dimensional logistic regression models, we define…

Methodology · Statistics 2022-10-06 Rong Ma , Zijian Guo , T. Tony Cai , Hongzhe Li

Clustering is a central approach for unsupervised learning. After clustering is applied, the most fundamental analysis is to quantitatively compare clusterings. Such comparisons are crucial for the evaluation of clustering methods as well…

Machine Learning · Statistics 2017-10-03 Alexander J Gates , Yong-Yeol Ahn