English
Related papers

Related papers: HEDE: Heritability estimation in high dimensions b…

200 papers

Decision trees and random forest remain highly competitive for classification on medium-sized, standard datasets due to their robustness, minimal preprocessing requirements, and interpretability. However, a single tree suffers from high…

Machine Learning · Statistics 2025-12-02 Cencheng Shen , Yuexiao Dong , Carey E. Priebe

This paper proposes a new method for estimating high-dimensional binary choice models. We consider a semiparametric model that places no distributional assumptions on the error term, allows for heteroskedastic errors, and permits endogenous…

Econometrics · Economics 2025-07-15 Fu Ouyang , Thomas Tao Yang

In Genome-Wide Association Studies (GWAS), heritability is defined as the fraction of variance of an outcome explained by a large number of genetic predictors in a high-dimensional polygenic linear model. This work studies the asymptotic…

Statistics Theory · Mathematics 2025-02-27 David Azriel , Samuel Davenport , Armin Schwartzman

Focusing on a high dimensional linear model $y = X\beta + \epsilon$ with dependent, non-stationary, and heteroskedastic errors, this paper applies the debiased and threshold ridge regression method that gives a consistent estimator for…

Statistics Theory · Mathematics 2021-10-27 Yunyi Zhang , Dimitris N. Politis

Analyzing data from multiple sources offers valuable opportunities to improve the estimation efficiency of causal estimands. However, this analysis also poses many challenges due to population heterogeneity and data privacy constraints.…

Methodology · Statistics 2025-10-23 Rong Zhao , Jason Falvey , Xu Shi , Vernon M. Chinchilli , Chixiang Chen

We study methods for simultaneous analysis of many noisy and biased estimates, each paired with an even noisier estimate of its own bias. The analyst's goal is to construct short calibrated intervals for each parameter. The standard…

Methodology · Statistics 2026-05-11 Wanyi Ling , Sida Li , Junming Guan , Nikolaos Ignatiadis

Electronic health records (EHRs) linked with familial relationship data offer a unique opportunity to investigate the genetic architecture of complex phenotypes at scale. However, existing heritability and coheritability estimation methods…

Methodology · Statistics 2025-11-12 Yinjun Zhao , Nicholas Tatonetti , Yuanjia Wang

In this study we propose a hybrid estimation of distribution algorithm (HEDA) to solve the joint stratification and sample allocation problem. This is a complex problem in which each the quality of each stratification from the set of all…

Methodology · Statistics 2022-01-12 Mervyn O'Luing , Steven Prestwich , S. Armagan Tarim

The method of instrumental variables provides a fundamental and practical tool for causal inference in many empirical studies where unmeasured confounding between the treatments and the outcome is present. Modern data such as the genetical…

Methodology · Statistics 2022-10-28 Ziang Niu , Yuwen Gu , Wei Li

The simultaneous estimation of many parameters based on data collected from corresponding studies is a key research problem that has received renewed attention in the high-dimensional setting. Many practical situations involve heterogeneous…

Methodology · Statistics 2026-03-26 Trambak Banerjee , Luella J. Fu , Gareth M. James , Gourab Mukherjee , Wenguang Sun

Image classification technology and performance based on Deep Learning have already achieved high standards. Nevertheless, many efforts have conducted to improve the stability of classification via ensembling. However, the existing ensemble…

Computer Vision and Pattern Recognition · Computer Science 2021-04-12 YeongHyeon Park , JoonSung Lee , Wonseok Park

Heterogeneous treatment effects (HTE) based on patients' genetic or clinical factors are of significant interest to precision medicine. Simultaneously modeling HTE and corresponding main effects for randomized clinical trials with…

Machine Learning · Statistics 2023-02-06 Heng Chen , Michael L. LeBlanc , James Y. Dai

We tackle modelling and inference for variable selection in regression problems with many predictors and many responses. We focus on detecting hotspots, i.e., predictors associated with several responses. Such a task is critical in…

In observational studies, accurately characterizing variance is critical for sample size determination, yet unaccounted-for variability from propensity score estimation and the resulting weights limit the accuracy of standard variance…

Methodology · Statistics 2026-04-24 Taekwon Hong , Daeyoung Lim , Woojung Bae , Yong Ma

We investigate statistical properties of a likelihood approach to nonparametric estimation of a singular distribution using deep generative models. More specifically, a deep generative model is used to model high-dimensional data that are…

Machine Learning · Statistics 2023-03-29 Minwoo Chae , Dongha Kim , Yongdai Kim , Lizhen Lin

High-dimensional dense embeddings have become central to modern Information Retrieval, but many dimensions are noisy or redundant. Recently proposed DIME (Dimension IMportance Estimation), provides query-dependent scores to identify…

Information Retrieval · Computer Science 2026-04-13 Giulio D'Erasmo , Cesare Campagnano , Antonio Mallia , Pierpaolo Brutti , Nicola Tonellotto , Fabrizio Silvestri

We consider large-scale studies in which it is of interest to test a very large number of hypotheses, and then to estimate the effect sizes corresponding to the rejected hypotheses. For instance, this setting arises in the analysis of gene…

Methodology · Statistics 2015-03-31 Kean Ming Tan , Noah Simon , Daniela Witten

The discovery of disease biomarkers from gene expression data has been greatly advanced by feature selection (FS) methods, especially using ensemble FS (EFS) strategies with perturbation at the data level (i.e., homogeneous, Hom-EFS) or…

Machine Learning · Computer Science 2021-08-03 Felipe Colombelli , Thayne Woycinck Kowalski , Mariana Recamonde-Mendoza

In this paper, we propose a new method to perform data augmentation in a reliable way in the High Dimensional Low Sample Size (HDLSS) setting using a geometry-based variational autoencoder. Our approach combines a proper latent space…

Machine Learning · Statistics 2023-01-18 Clément Chadebec , Elina Thibeau-Sutre , Ninon Burgos , Stéphanie Allassonnière

High-dimensional, low sample-size (HDLSS) data problems have been a topic of immense importance for the last couple of decades. There is a vast literature that proposed a wide variety of approaches to deal with this situation, among which…

Methodology · Statistics 2021-07-09 Kaixu Yang , Tapabrata Maiti