English
Related papers

Related papers: Agglomerative Hierarchical Clustering for Selectin…

200 papers

Cluster analysis has proved to be an invaluable tool for the exploratory and unsupervised analysis of high dimensional datasets. Among methods for clustering, hierarchical approaches have enjoyed substantial popularity in genomics and other…

Methodology · Statistics 2014-11-20 Patrick K. Kimes , Yufeng Liu , D. Neil Hayes , J. S. Marron

Matrix valued data has become increasingly prevalent in many applications. Most of the existing clustering methods for this type of data are tailored to the mean model and do not account for the dependence structure of the features, which…

Machine Learning · Statistics 2023-12-07 Inbeom Lee , Siyi Deng , Yang Ning

In various practical situations, forecasting of aggregate values rather than individual ones is often our main focus. For instance, electricity companies are interested in forecasting the total electricity demand in a specific region to…

Methodology · Statistics 2025-08-22 Kei Hirose , Hidetoshi Matsui , Hiroki Masuda

VARCLUST algorithm is proposed for clustering variables under the assumption that variables in a given cluster are linear combinations of a small number of hidden latent variables, corrupted by the random noise. The entire clustering task…

Clustering of mixed-type datasets can be a particularly challenging task as it requires taking into account the associations between variables with different level of measurement, i.e., nominal, ordinal and/or interval. In some cases,…

Methodology · Statistics 2022-04-22 Odysseas Moschidis , Angelos Markos , Theodore Chadjipadelis

Studies investigating the causal effects of spatially varying exposures on outcomes often rely on observational and spatially indexed data. A prevalent challenge is unmeasured spatial confounding, where an unobserved spatially varying…

Methodology · Statistics 2025-11-19 Sophie M. Woodward , Mauricio Tec , Francesca Dominici

We develop a highly scalable optimization method called "hierarchical group-thresholding" for solving a multi-task regression model with complex structured sparsity constraints on both input and output spaces. Despite the recent emergence…

Machine Learning · Statistics 2012-08-16 Seunghak Lee , Eric P. Xing

Instrumental variables (IV) regression is widely used to estimate causal treatment effects in settings where receipt of treatment is not fully random, but there exists an instrument that generates exogenous variation in treatment exposure.…

Econometrics · Economics 2021-08-10 Stephen Coussens , Jann Spiess

Clustering is the technique to partition data according to their characteristics. Data that are similar in nature belong to the same cluster [1]. There are two types of evaluation methods to evaluate clustering quality. One is an external…

Machine Learning · Computer Science 2024-09-05 Anupriya Vysala , Joseph Gomes

This paper proposes semi-instrumental variables (semi-IVs) as an alternative to instrumental variables (IVs) to identify the causal effect of a binary (or discrete) endogenous treatment. A semi-IV is a less restrictive form of instrument:…

Econometrics · Economics 2025-09-23 Christophe Bruneel-Zupanc

Conventional cluster-robust inference can be invalid when data contain clusters of unignorably large size. We formalize this issue by deriving a necessary and sufficient condition for its validity, and show that this condition is frequently…

Econometrics · Economics 2025-10-07 Harold D. Chiang , Yuya Sasaki , Yulong Wang

In this paper, we discuss causal inference on the efficacy of a treatment or medication on a time-to-event outcome with competing risks. Although the treatment group can be randomized, there can be confoundings between the compliance and…

Methodology · Statistics 2016-12-06 Cheng Zheng , Ran Dai , Parameswaran Hari , Mei-Jie Zhang

When fitting statistical models, some predictors are often found to be correlated with each other, and functioning together. Many group variable selection methods are developed to select the groups of predictors that are closely related to…

Methodology · Statistics 2021-03-25 Zhiyuan Li

We show how clustering standard errors in one or more dimensions can be justified in M-estimation when there is sampling or assignment uncertainty. Since existing procedures for variance estimation are either conservative or invalid, we…

Econometrics · Economics 2024-11-21 Ruonan Xu , Luther Yap

Many treatment variables used in empirical applications nest multiple unobserved versions of a treatment. I show that instrumental variable (IV) estimands for the effect of a composite treatment are IV-specific weighted averages of effects…

General Economics · Economics 2022-11-24 Clint Harris

Hierarchical categorical variables often exhibit many levels (high granularity) and many classes within each level (high dimensionality). This may cause overfitting and estimation issues when including such covariates in a predictive model.…

Methodology · Statistics 2024-08-20 Paul Wilsens , Katrien Antonio , Gerda Claeskens

We propose a weak-identification-robust test for linear instrumental variable (IV) regressions with high-dimensional instruments, whose number is allowed to exceed the sample size. In addition, our test is robust to general error…

Econometrics · Economics 2025-07-01 Qu Feng , Sombut Jaidee , Wenjie Wang

We consider the challenging problem of estimating causal effects from purely observational data in the bi-directional Mendelian randomization (MR), where some invalid instruments, as well as unmeasured confounding, usually exist. To address…

Methodology · Statistics 2024-07-15 Feng Xie , Zhen Yao , Lin Xie , Yan Zeng , Zhi Geng

This paper focuses on variable selection for a partially linear single-index varying-coefficient model. A regularized variable selection procedure by combining basis function approximations with SCAD penalty is proposed. It can…

Statistics Theory · Mathematics 2024-12-19 Lijuan Han , Liugen Xue , Junshan Xie

We consider the problem of model-based clustering in the presence of many correlated, mixed continuous and discrete variables, some of which may have missing values. Discrete variables are treated with a latent continuous variable approach…