English
Related papers

Related papers: Adjusting for informative cluster size in pseudo-v…

200 papers

Counterfactual explanations for black-box models aim to pr ovide insight into an algorithmic decision to its recipient. For a binary classification problem an individual counterfactual details which features might be changed for the model…

Machine Learning · Statistics 2025-05-29 James M. Adams , Gesine Reinert , Lukasz Szpruch , Carsten Maple , Andrew Elliott

The COVID-19 pandemic has taken the world by storm with its high infection rate. Investigating its geographical disparities has paramount interest in order to gauge its relationships with political decisions, economic indicators, or mental…

Applications · Statistics 2023-12-29 Amay SM Cheam , Marc Fredette , Matthieu Marbac , Fabien Navarro

Instrumental variable methods are widely used in medical and social science research to draw causal conclusions when the treatment and outcome are confounded by unmeasured confounding variables. One important feature of such studies is that…

Methodology · Statistics 2021-05-25 Bo Zhang , Siyu Heng , Emily J. MacKay , Ting Ye

Having reliable estimates of the occurrence rates of extreme events is highly important for insurance companies, government agencies and the general public. The rarity of an extreme event is typically expressed through its return period,…

Methodology · Statistics 2019-10-08 Ross Towe , Jonathan Tawn , Emma Eastoe , Rob Lamb

Consensus clustering has been widely used in bioinformatics and other applications to improve the accuracy, stability and reliability of clustering results. This approach ensembles cluster co-occurrences from multiple clustering runs on…

Machine Learning · Statistics 2023-01-11 Luqin Gan , Genevera I. Allen

Propensity score weighting is a tool for causal inference to adjust for measured confounders. Survey data are often collected under complex sampling designs such as multistage cluster sampling, which presents challenges for propensity score…

Methodology · Statistics 2016-07-27 Shu Yang

Temporal data, obtained in the setting where it is only possible to observe one time point per experiment, is widely used in different research fields, yet remains insufficiently addressed from the statistical point of view. Such data often…

Methodology · Statistics 2025-03-10 Polina Arsenteva , Mohamed Amine Benadjaoud , Hervé Cardot

We propose a method for high dimensional multivariate regression that is robust to random error distributions that are heavy-tailed or contain outliers, while preserving estimation accuracy in normal random error distributions. We extend…

Methodology · Statistics 2025-03-05 Mayu Hiraishi , Kensuke Tanioka , Hiroshi Yadohisa

Self-supervised learning approaches provide a promising direction for clustering multivariate time-series data. However, real-world time-series data often include missing values, and the existing approaches require imputing missing values…

Machine Learning · Computer Science 2023-05-30 Hamid Ghaderi , Brandon Foreman , Amin Nayebi , Sindhu Tipirneni , Chandan K. Reddy , Vignesh Subbian

In many modern statistical problems, the limited available data must be used both to develop the hypotheses to test, and to test these hypotheses-that is, both for exploratory and confirmatory data analysis. Reusing the same dataset for…

Methodology · Statistics 2023-07-24 Youngjoo Yun , Rina Foygel Barber

We introduce usage of a reduction property of penalty-based formulation of pseudo-Boolean polynomials as a mechanism for invariant dimensionality reduction in cluster analysis processes. In our experiments, we show that multidimensional…

Information Retrieval · Computer Science 2023-08-31 Tendai Mapungwana Chikake , Boris Goldengorin

Representative risk estimation is fundamental to clinical decision-making. However, risks are often estimated from non-representative epidemiologic studies, which usually underrepresent minorities. "Model-based" methods use population…

Methodology · Statistics 2023-04-12 Lingxiao Wang , Yan Li , Barry I. Graubard , Hormuzd A. Katki

To estimate casual treatment effects, we propose a new matching approach based on the reduced covariates obtained from sufficient dimension reduction. Compared to the original covariates and the propensity score, which are commonly used for…

Methodology · Statistics 2017-02-03 Wei Luo , Yeying Zhu

Personalization is very powerful in improving the effectiveness of health interventions. Reinforcement learning (RL) algorithms are suitable for learning these tailored interventions from sequential data collected about individuals.…

Artificial Intelligence · Computer Science 2020-05-22 Ali el Hassouni , Mark Hoogendoorn , Martijn van Otterlo , A. E. Eiben , Vesa Muhonen , Eduardo Barbaro

This paper describes an approach to simultaneously identify clusters and estimate cluster-specific regression parameters from the given data. Such an approach can be useful in learning the relationship between input and output when the…

Statistical Finance · Quantitative Finance 2024-01-02 Udai Nagpal , Krishan Nagpal

In applications where the study data are collected within cluster units (e.g., patients within transplant centers), it is often of interest to estimate and perform inference on the treatment effects of the cluster units. However, it is…

Applications · Statistics 2023-05-11 Nicholas Hartman , Kevin He

Importance sampling (IS) is a common reweighting strategy for off-policy prediction in reinforcement learning. While it is consistent and unbiased, it can result in high variance updates to the weights for the value function. In this work,…

Machine Learning · Computer Science 2019-11-15 Matthew Schlegel , Wesley Chung , Daniel Graves , Jian Qian , Martha White

Individualized treatment regimes (ITRs) aim to improve clinical outcomes by assigning treatment based on patient-specific characteristics. However, existing methods often struggle with high-dimensional covariates, limiting accuracy,…

Machine Learning · Statistics 2026-01-13 Sungtaek Son , Eardi Lila , Kwun Chuen Gary Chan

We propose In-Context Clustering (ICC), a flexible LLM-based procedure for clustering data from diverse distributions. Unlike traditional clustering algorithms constrained by predefined similarity measures, ICC flexibly captures complex…

Machine Learning · Computer Science 2025-10-10 Ying Wang , Mengye Ren , Andrew Gordon Wilson

Predicting time-to-event outcomes in large databases can be a challenging but important task. One example of this is in predicting the time to a clinical outcome for patients in intensive care units (ICUs), which helps to support critical…

Computation · Statistics 2019-08-06 Yingying Xu , Joon Lee , Joel A. Dubin