English
Related papers

Related papers: Sufficient Dimension Reduction for Populations wit…

200 papers

Estimation of heterogeneous treatment effects is an active area of research. Most of the existing methods, however, focus on estimating the conditional average treatment effects of a single, binary treatment given a set of pre-treatment…

Methodology · Statistics 2025-05-30 Max Goplerud , Kosuke Imai , Nicole E. Pashley

In this study, we address the challenge of survival analysis within heterogeneous patient populations, where traditional reliance on a single regression model such as the Cox proportional hazards (Cox PH) model often falls short.…

Methodology · Statistics 2025-04-15 Ahmad Talafha

This paper explores the homogeneity of coefficients in high-dimensional regression, which extends the sparsity concept and is more general and suitable for many applications. Homogeneity arises when one expects regression coefficients…

Methodology · Statistics 2013-04-01 Tracy Ke , Jianqing Fan , Yichao Wu

Environmental hazards place certain individuals at disproportionately higher risks. As these hazards increasingly endanger human health, precise identification of the most vulnerable population subgroups is critical for public health.…

Machine Learning · Computer Science 2024-09-23 Jong Woo Nam , Eun Young Choi , Jennifer A. Ailshire , Yao-Yi Chiang

Risk modeling with EHR data is challenging due to a lack of direct observations on the disease outcome, and the high dimensionality of the candidate predictors. In this paper, we develop a surrogate assisted semi-supervised-learning (SAS)…

Statistics Theory · Mathematics 2021-05-05 Jue Hou , Zijian Guo , Tianxi Cai

This study proposes a Transformer-based longitudinal modeling method to address challenges in clinical risk classification with heterogeneous Electronic Health Record (EHR) data, including irregular temporal patterns, large modality…

Machine Learning · Computer Science 2025-11-07 Anzhuo Xie , Wei-Chen Chang

There is increasing interest in modeling high-dimensional longitudinal outcomes in applications such as developmental neuroimaging research. Growth curve model offers a useful tool to capture both the mean growth pattern across individuals,…

Methodology · Statistics 2023-05-26 Lu Wang , Xiang Lyu , Zhengwu Zhang , Lexin Li

Data heterogeneity plays a pivotal role in determining the performance of machine learning (ML) systems. Traditional algorithms, which are typically designed to optimize average performance, often overlook the intrinsic diversity within…

Machine Learning · Computer Science 2025-06-03 Jiashuo Liu , Peng Cui

Nowadays, massive datasets are typically dispersed across multiple locations, encountering dual challenges of high dimensionality and huge sample size. Therefore, it is necessary to explore sufficient dimension reduction (SDR) methods for…

Methodology · Statistics 2025-09-16 Hongying Li , Minyi Zhu , Yaqi Cao , Xinyi Xu

In microbiome and genomic studies, the regression of compositional data has been a crucial tool for identifying microbial taxa or genes that are associated with clinical phenotypes. To account for the variation in sequencing depth, the…

Methodology · Statistics 2021-03-11 Pixu Shi , Yuchen Zhou , Anru R. Zhang

How to estimate heterogeneity, e.g. the effect of some variable differing across observations, is a key question in political science. Methods for doing so make simplifying assumptions about the underlying nature of the heterogeneity to…

Methodology · Statistics 2021-03-31 Max Goplerud

Problem: How can we design patient cost prediction models that effectively address the challenges of heterogeneity in administrative claims (AC) data to ensure accurate, fair, and generalizable predictions, especially for high-need (HN)…

Machine Learning · Computer Science 2025-02-19 Mohammad Amin Morid , Olivia R. Liu Sheng

Many data sets contain an inherent multilevel structure, for example, because of repeated measurements of the same observational units. Taking this structure into account is critical for the accuracy and calibration of any statistical…

Methodology · Statistics 2020-05-07 Topi Paananen , Alejandro Catalina , Paul-Christian Bürkner , Aki Vehtari

In many biomedical problems, data are often heterogeneous, with samples spanning multiple patient subgroups, where different subgroups may have different disease subtypes, stages, or other medical contexts. These subgroups may be related,…

Methodology · Statistics 2022-11-30 Zihan Li , Ziye Luo , Yifan Sun

Repeated-measure designs allow comparisons within a group as well as between groups, and are commonly referred to as split-plot designs. While originating in agricultural experiments, they are now widely used in medical research,…

Computation · Statistics 2025-12-22 Paavo Sattler , Nils Hichert

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

Methodology · Statistics 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu

The estimation of causal treatment effects from observational data is a fundamental problem in causal inference. To avoid bias, the effect estimator must control for all confounders. Hence practitioners often collect data for as many…

Machine Learning · Statistics 2020-11-05 Kristjan Greenewald , Dmitriy Katz-Rogozhnikov , Karthik Shanmugam

Heterogeneity is a dominant factor in the behaviour of many biological processes. Despite this, it is common for mathematical and statistical analyses to ignore biological heterogeneity as a source of variability in experimental data.…

High-dimensional compositional data, such as those from human microbiome studies, pose unique statistical challenges due to the simplex constraint and excess zeros. While dimension reduction is indispensable for analyzing such data,…

Methodology · Statistics 2025-09-09 Junyoung Park , Cheolwoo Park , Jeongyoun Ahn

In this paper, we address the problem of predicting a response variable in the context of both, spatially correlated and high-dimensional data. To reduce the dimensionality of the predictor variables, we apply the sufficient dimension…

Methodology · Statistics 2025-02-06 Liliana Forzani , Rodrigo García Arancibia , Antonella Gieco , Pamela Llop , Anne Yao