中文
相关论文

相关论文: Agglomerative Hierarchical Clustering for Selectin…

200 篇论文

We survey agglomerative hierarchical clustering algorithms and discuss efficient implementations that are available in R and other software environments. We look at hierarchical self-organizing maps, and mixture models. We review grid-based…

信息检索 · 计算机科学 2011-05-03 Fionn Murtagh , Pedro Contreras

Experiments studying get-out-the-vote (GOTV) efforts estimate the causal effect of various mobilization efforts on voter turnout. However, there is often substantial noncompliance in these studies. A usual approach is to use an instrumental…

统计方法学 · 统计学 2024-07-02 Nicole E. Pashley , Luke Keele , Luke W. Miratrix

Traditional instrumental variable (IV) estimators face a fundamental constraint: they can only accommodate as many endogenous treatment variables as available instruments. This limitation becomes particularly challenging in settings where…

机器学习 · 计算机科学 2025-06-25 Shiangyi Lin , Hui Lan , Vasilis Syrgkanis

We consider the problem of testing for differences in group-specific slopes between the selected groups in panel data identified via k-means clustering. In this setting, the classical Wald-type test statistic is problematic because it…

统计方法学 · 统计学 2025-11-07 Chuang Wan , Jiajun Sun , Xingbai Xu

Variable selection comprises an important step in many modern statistical inference procedures. In the regression setting, when estimators cannot shrink irrelevant signals to zero, covariates without relationships to the response often…

统计理论 · 数学 2025-03-28 Ka Long Keith Ho , Hien Duy Nguyen

In cluster analysis, it can be useful to interpret the partition built from the data in the light of external categorical variables which were not directly involved to cluster the data. An approach is proposed in the model-based clustering…

The prevailing IR-threshold paradigm posits a positive correlation between imbalance ratio (IR) and oversampling effectiveness, yet this assumption remains empirically unsubstantiated through controlled experimentation. We conducted 12…

机器学习 · 计算机科学 2026-04-07 Yuwen Jiang , Songyun Ye

The instrumental-variables (IV) setting is standard for partial identification of causal effects when unobserved confounding makes point identification impossible. Existing approaches face methodological bottlenecks: closed-form bound…

机器学习 · 计算机科学 2026-05-14 Vahid Balazadeh , Hamidreza Kamkari , Medha Barath , Ricardo Silva , Rahul G. Krishnan

The mixture models have become widely used in clustering, given its probabilistic framework in which its based, however, for modern databases that are characterized by their large size, these models behave disappointingly in setting out the…

机器学习 · 统计学 2017-02-01 Abdelghafour Talibi , Boujemâa Achchab , Rafik Lasri

The applicability of agglomerative clustering, for inferring both hierarchical and flat clustering, is limited by its scalability. Existing scalable hierarchical clustering methods sacrifice quality for speed and often lead to over-merging…

This paper uses model symmetries in the instrumental variable (IV) regression to derive an invariant test for the causal structural parameter. Contrary to popular belief, we show that there exist model symmetries when equation errors are…

统计理论 · 数学 2021-09-02 Marcelo J. Moreira , Mahrad Sharifvaghefi , Geert Ridder

Ensemble Learning methods combine multiple algorithms performing the same task to build a group with superior quality. These systems are well adapted to the distributed setup, where each peer or machine of the network hosts one algorithm…

机器学习 · 计算机科学 2021-10-19 Gaëlle Candel , David Naccache

This paper studies the challenging problem of estimating causal effects from observational data, in the presence of unobserved confounders. The two-stage least square (TSLS) method and its variants with a standard instrumental variable (IV)…

机器学习 · 计算机科学 2023-10-04 Debo Cheng , Ziqi Xu , Jiuyong Li , Lin Liu , Jixue Liu , Thuc Duy Le

Crowdsourcing has become an effective and popular tool for human-powered computation to label large datasets. Since the workers can be unreliable, it is common in crowdsourcing to assign multiple workers to one task, and to aggregate the…

机器学习 · 统计学 2014-11-18 Hongwei Li , Bin Yu

To reach human level intelligence, learning algorithms need to incorporate causal reasoning. But identifying causality, and particularly counterfactual reasoning, remains elusive. In this paper, we make progress on counterfactual inference…

机器学习 · 统计学 2026-03-31 Marc Braun , Jose M. Peña , Adel Daoud

Multiple imputation has become one of the standard methods in drawing inferences in many incomplete data applications. Applications of multiple imputation in relatively more complex settings, such as high-dimensional clustered data, require…

统计方法学 · 统计学 2025-04-08 Qiushuang Li , Recai Yucel

Cluster analysis is a popular unsupervised learning tool used in many disciplines to identify heterogeneous sub-populations within a sample. However, validating cluster analysis results and determining the number of clusters in a data set…

机器学习 · 统计学 2024-04-26 Ali Turfah , Xiaoquan Wen

Classification of cluster variables in cluster algebras (in particular, Grassmannian cluster algebras) is an important problem, which has direct application to computations of scattering amplitudes in physics. In this paper, we apply the…

高能物理 - 理论 · 物理学 2026-02-16 Man-Wai Cheung , Pierre-Philippe Dechant , Yang-Hui He , Elli Heyes , Edward Hirst , Jian-Rong Li

Instrumental variable regression is a common approach for causal inference in the presence of unobserved confounding. However, identifying valid instruments is often difficult in practice. In this paper, we propose a novel method based on…

统计方法学 · 统计学 2026-01-22 Gregor Steiner , Jeremie Houssineau , Mark F. J. Steel

This article investigates uncertainty quantification of the generalized linear lasso~(GLL), a popular variable selection method in high-dimensional regression settings. In many fields of study, researchers use data-driven methods to select…

统计理论 · 数学 2023-07-11 Quentin Duchemin , Yohann de Castro