English
Related papers

Related papers: Leverage, Influence, and the Jackknife in Clustere…

200 papers

Stepped wedge designs (SWDs) are increasingly used to evaluate longitudinal cluster-level interventions but pose substantial challenges for valid inference. Because crossover times are randomized, intervention effects are intrinsically…

Methodology · Statistics 2026-05-12 Fan Xia , K. C. Gary Chan , Emily Voldal , Avi Kenny , Patrick J. Heagerty , James P. Hughes

We present a machine learning framework and a new test bed for data mining from the Slurm Workload Manager for high-performance computing (HPC) clusters. The focus was to find a method for selecting features to support decisions: helping…

Machine Learning · Computer Science 2020-12-16 Adedolapo Okanlawon , Huichen Yang , Avishek Bose , William Hsu , Dan Andresen , Mohammed Tanash

Survival analysis aims to predict the timing of future events across various fields, from medical outcomes to customer churn. However, the integration of clustering into survival analysis, particularly for precision medicine, remains…

Machine Learning · Computer Science 2024-05-28 Gabriel Buginga , Edmundo de Souza e Silva

Due to its ability to combine multiple base clusterings into a probably better and more robust clustering, the ensemble clustering technique has been attracting increasing attention in recent years. Despite the significant success, one…

Machine Learning · Computer Science 2020-01-01 Dong Huang , Chang-Dong Wang , Jian-Huang Lai

We propose a new approach for sparse regression and marginal testing, for data with correlated features. Our procedure first clusters the features, and then chooses as the cluster prototype the most informative feature in that cluster. Then…

Methodology · Statistics 2015-03-16 Stephen Reid , Robert Tibshirani

The manuscript discusses how to incorporate random effects for quantile regression models for clustered data with focus on settings with many but small clusters. The paper has three contributions: (i) documenting that existing methods may…

Methodology · Statistics 2022-02-24 Maria Laura Battagliola , Helle Sørensen , Anders Tolver , Ana-Maria Staicu

A general jackknife estimator for the asymptotic covariance of moment estimators is considered in the case when the sample is taken from a mixture with varying concentrations of components. Consistency of the estimator is demonstrated. A…

Statistics Theory · Mathematics 2019-12-18 Rostyslav Maiboroda , Olena Sugakova

In economic development, there are often regions that share similar economic characteristics, and economic models on such regions tend to have similar covariate effects. In this paper, we propose a Bayesian clustered regression for…

Econometrics · Economics 2020-06-30 Zhihua Ma , Yishu Xue , Guanyu Hu

We introduce a general semiparametric clusterwise elliptical distribution to assess how latent cluster structure shapes continuous outcomes. Using a subjectwise representation, we first estimate cluster-specific mean vectors and a…

Methodology · Statistics 2026-04-10 Jen-Chieh Teng , Sheng-Hsin Fan , Chin-Tsang Chiang , Ming-Yueh Huang , Alvin Lim

Stepped wedge cluster randomized trials (SW-CRTs) have become increasingly popular and are used for a variety of interventions and outcomes, often chosen for their feasibility advantages. SW-CRTs must account for time trends in the outcome…

Methodology · Statistics 2024-07-16 Lee Kennedy-Shaffer , Victor De Gruttola , Marc Lipsitch

We develop a concept of weak identification in linear IV models in which the number of instruments can grow at the same rate or slower than the sample size. We propose a jackknifed version of the classical weak identification-robust…

Econometrics · Economics 2021-10-06 Anna Mikusheva , Liyang Sun

We propose a novel clustering pipeline to detect and characterize influence campaigns from documents. This approach clusters parts of document, detects clusters that likely reflect an influence campaign, and then identifies documents linked…

Computation and Language · Computer Science 2024-04-30 Zhengxiang Wang , Owen Rambow

Samples with a common mean but possibly different, ordered variances arise in various fields such as interlaboratory experiments, field studies or the analysis of sensor data. Estimators for the common mean under ordered variances typically…

Statistics Theory · Mathematics 2019-01-30 Ansgar Steland , Yuan-Tsung Chang

Clustering is a fundamental tool in unsupervised learning, used to group objects by distinguishing between similar and dissimilar features of a given data set. One of the most common clustering algorithms is k-means. Unfortunately, when…

Machine Learning · Statistics 2021-08-17 Olga Dorabiala , J. Nathan Kutz , Aleksandr Aravkin

We introduce a new method for performing clustering with the aim of fitting clusters with different scatters and weights. It is designed by allowing to handle a proportion $\alpha$ of contaminating data to guarantee the robustness of the…

Statistics Theory · Mathematics 2008-12-18 Luis A. García-Escudero , Alfonso Gordaliza , Carlos Matrán , Agustin Mayo-Iscar

When scholars suspect units are dependent on each other within clusters but independent of each other across clusters, they employ cluster-robust standard errors (CRSEs). Nevertheless, what to cluster over is sometimes unknown. For…

Methodology · Statistics 2025-11-12 Kentaro Fukumoto

Clustering is widely used for unsupervised structure discovery, yet it offers limited insight into how reliable each individual assignment is. Diagnostics, such as convergence behavior or objective values, may reflect global quality, but…

Machine Learning · Computer Science 2026-05-15 Aggelos Semoglou , John Pavlopoulos

We develop tools for selective inference in the setting of group sparsity, including the construction of confidence intervals and p-values for testing selected groups of variables. Our main technical result gives the precise distribution of…

Methodology · Statistics 2016-07-28 Fan Yang , Rina Foygel Barber , Prateek Jain , John Lafferty

We introduce a novel class of Bayesian mixtures for normal linear regression models which incorporates a further Gaussian random component for the distribution of the predictor variables. The proposed cluster-weighted model aims to…

Methodology · Statistics 2026-05-26 Panagiotis Papastamoulis , Konstantinos Perrakis

It has become standard for empirical studies to conduct inference robust to cluster dependence and heterogeneity. With a small number of clusters, the normal approximation for the $t$-statistics of regression coefficients may be poor. This…

Econometrics · Economics 2026-03-27 Bulat Gafarov , Takuya Ura