English
Related papers

Related papers: A Criterion for Aggregation Error for Multivariate…

200 papers

Many modern datasets consist of multiple related matrices measured on a common set of units, where the goal is to recover the shared low-dimensional subspace. While the Angle-based Joint and Individual Variation Explained (AJIVE) framework…

Statistics Theory · Mathematics 2025-12-03 Jingyang Li , Zhongyuan Lyu

The deployment of machine learning classifiers in high-stakes domains requires well-calibrated confidence scores for model predictions. In this paper we introduce the notion of variable-based calibration to characterize calibration…

Machine Learning · Computer Science 2023-04-07 Markelle Kelly , Padhraic Smyth

Clustering is widely used for unsupervised structure discovery, yet it offers limited insight into how reliable each individual assignment is. Diagnostics, such as convergence behavior or objective values, may reflect global quality, but…

Machine Learning · Computer Science 2026-05-15 Aggelos Semoglou , John Pavlopoulos

In the context of the Classification and Regression Trees (CART) algorithm, the efficient splitting of categorical features using standard criteria like GINI and Entropy is well-established. However, using the Mean Absolute Error (MAE)…

Machine Learning · Computer Science 2025-11-12 Peng Yu , Yike Chen , Chao Xu , Albert Bifet , Jesse Read

Causal effect estimation (CEE) provides a crucial tool for predicting the unobserved counterfactual outcome for an entity. As CEE relaxes the requirement for ``perfect'' counterfactual samples (e.g., patients with identical attributes and…

Machine Learning · Computer Science 2024-11-19 Hechuan Wen , Tong Chen , Guanhua Ye , Li Kheng Chai , Shazia Sadiq , Hongzhi Yin

The objective of this study is to investigate spatial structures of error in the assessment of continuous raster data. The use of conventional diagnostics of error often overlooks the possible spatial variation in error because such…

Applications · Statistics 2025-10-15 Narumasa Tsutsumida , Pedro Rodríguez-Veiga , Paul Harris , Heiko Balzter , Alexis Comber

We decompose the Kullback--Leibler generalization error (GE) -- the expected KL divergence from the data distribution to the trained model -- of unsupervised learning into three non-negative components: model error, data bias, and variance.…

Machine Learning · Statistics 2026-04-15 Gilhan Kim

Multi-view clustering (MVC) aims to integrate complementary information from multiple views to enhance clustering performance. Late Fusion Multi-View Clustering (LFMVC) has shown promise by synthesizing diverse clustering results into a…

Machine Learning · Computer Science 2024-12-25 Liang Du , Henghui Jiang , Xiaodong Li , Yiqing Guo , Yan Chen , Feijiang Li , Peng Zhou , Yuhua Qian

Studying unified model averaging estimation for situations with complicated data structures, we propose a novel model averaging method based on cross-validation (MACV). MACV unifies a large class of new and existing model averaging…

Methodology · Statistics 2024-12-16 Dalei Yu , Xinyu Zhang , Hua Liang

Originally introduced as a neural network for ensemble learning, mixture of experts (MoE) has recently become a fundamental building block of highly successful modern deep neural networks for heterogeneous data analysis in several…

Machine Learning · Statistics 2024-02-12 Huy Nguyen , TrungTin Nguyen , Khai Nguyen , Nhat Ho

In modern randomized experiments, large-scale data collection increasingly yields rich baseline covariates and auxiliary information from multiple sources. Such information offers opportunities for more precise treatment effect estimation,…

Methodology · Statistics 2026-03-10 Wei Ma , Zeqi Wu , Zheng Zhang

Loss functions play a central role in supervised classification. Cross-entropy (CE) is widely used, whereas the mean absolute error (MAE) loss can offer robustness but is difficult to optimize. Interpolating between the CE and MAE losses,…

Machine Learning · Statistics 2026-04-29 Kartheek Bondugula , Santiago Mazuelas , Aritz Pérez , Anqi Liu

Linear regression model (LRM) based on mean square error (MSE) criterion is widely used in Granger causality analysis (GCA), which is the most commonly used method to detect the causality between a pair of time series. However, when signals…

Methodology · Statistics 2019-02-20 Badong Chen , Rongjin Ma , Siyu Yu , Shaoyi Du , Jing Qin

The model uncertainty obtained by variational Bayesian inference with Monte Carlo dropout is prone to miscalibration. In this paper, different logit scaling methods are extended to dropout variational inference to recalibrate model…

Machine Learning · Computer Science 2020-06-23 Max-Heinrich Laves , Sontje Ihler , Karl-Philipp Kortmann , Tobias Ortmaier

This paper investigates the cross-correlations across multiple climate model errors. We build a Bayesian hierarchical model that accounts for the spatial dependence of individual models as well as cross-covariances across different climate…

Applications · Statistics 2012-03-02 Huiyan Sang , Mikyoung Jun , Jianhua Z. Huang

Multiplexed Assays of Variant Effect (MAVEs) have emerged as a powerful approach for interrogating thousands of genetic variants in a single experiment. The flexibility and widespread adoption of these techniques across diverse disciplines…

Recently, contrastive learning (CL) plays an important role in exploring complementary information for multi-view clustering (MVC) and has attracted increasing attention. Nevertheless, real-world multi-view data suffer from data…

Machine Learning · Computer Science 2025-12-29 Hongqing He , Jie Xu , Wenyuan Yang , Yonghua Zhu , Guoqiu Wen , Xiaofeng Zhu

In order to obtain morphological information of unlabeled galaxies, we present an unsupervised machine-learning (UML) method for morphological classification of galaxies, which can be summarized as two aspects: (1) the methodology of…

Astrophysics of Galaxies · Physics 2022-02-02 C. C. Zhou , Y. Z. Gu , G. W. Fang , Z. S. Lin

Large spatiotemporal demand datasets can prove intractable for location optimization problems, motivating the need to aggregate such data. However, demand aggregation introduces error which impacts the results of the location study. We…

Applications · Statistics 2019-10-14 Zachary T. Hornberger , Bruce A. Cox , Raymond R. Hill

Land use and land cover mapping from Earth Observation (EO) data is a critical tool for sustainable land and resource management. While advanced machine learning and deep learning algorithms excel at analyzing EO imagery data, they often…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Babak Ghassemi , Cassio Fraga-Dantas , Raffaele Gaetano , Dino Ienco , Omid Ghorbanzadeh , Emma Izquierdo-Verdiguier , Francesco Vuolo