English
Related papers

Related papers: A Criterion for Aggregation Error for Multivariate…

200 papers

In health and social sciences, it is critically important to identify subgroups of the study population where there is notable heterogeneity of treatment effects (HTE) with respect to the population average. Decision trees have been…

Methodology · Statistics 2024-05-28 Falco J. Bargagli-Stoffi , Riccardo Cadei , Kwonsang Lee , Francesca Dominici

Linear regression with measurement error in the covariates is a heavily studied topic, however, the statistics/econometrics literature is almost silent to estimating a multi-equation model with measurement error. This paper considers a…

Methodology · Statistics 2020-06-15 Georges Bresson , Anoop Chaturvedi , Mohammad Arshad Rahman , Shalabh

In weakly supervised learning, unbiased risk estimator(URE) is a powerful tool for training classifiers when training and test data are drawn from different distributions. Nevertheless, UREs lead to overfitting in many problem settings when…

Machine Learning · Computer Science 2020-08-25 Yu-Ting Chou , Gang Niu , Hsuan-Tien Lin , Masashi Sugiyama

Causal machine learning holds promise for estimating individual treatment effects from complex data. For successful real-world applications of machine learning methods, it is of paramount importance to obtain reliable insights into which…

Machine Learning · Computer Science 2026-05-22 Joseph Paillard , Angel Reyero Lobo , Vitaliy Kolodyazhniy , Bertrand Thirion , Denis A. Engemann

This paper introduces the Mixed Aggregate Preference Logit (MAPL, pronounced "maple'') model, a novel class of discrete choice models that leverages machine learning to model unobserved heterogeneity in discrete choice analysis. The…

Econometrics · Economics 2025-03-05 Connor R. Forsythe , Cristian Arteaga , John P. Helveston

It is valuable for any decision maker to know the impact of decisions (treatments) on average and for subgroups. The causal machine learning literature has recently provided tools for estimating group average treatment effects (GATE) to…

Econometrics · Economics 2025-01-10 Nora Bearth , Michael Lechner

Confidence calibration is essential for making large language models (LLMs) reliable, yet existing training-free methods have been primarily studied under single-answer question answering. In this paper, we show that these methods break…

Computation and Language · Computer Science 2026-02-10 Yuhan Wang , Shiyu Ni , Zhikai Ding , Zihang Zhan , Yuanzi Li , Keping Bi

We propose a variational autoencoder (VAE) approach for parameter estimation in nonlinear mixed-effects models based on ordinary differential equations (NLME-ODEs) using longitudinal data from multiple subjects. In moderate dimensions,…

Methodology · Statistics 2026-02-11 Zhe Li , Mélanie Prague , Rodolphe Thiébaut , Quentin Clairon

Recently, data-driven weather forecasting methods have received significant attention for surpassing the RMSE performance of traditional NWP (Numerical Weather Prediction)-based methods. However, data-driven models are tuned to minimize the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Doyi Kim , Minseok Seo , Yeji Choi

Popular (ensemble) Kalman filter data assimilation (DA) approaches assume that the errors in both the a priori estimate of the state and those in the observations are Gaussian. For constrained variables, e.g. sea ice concentration or…

Machine Learning · Computer Science 2025-02-19 Ivo Pasmans , Yumeng Chen , Tobias Sebastian Finn , Marc Bocquet , Alberto Carrassi

A major effort in modern high-dimensional statistics has been devoted to the analysis of linear predictors trained on nonlinear feature embeddings via empirical risk minimization (ERM). Gaussian equivalence theory (GET) has emerged as a…

Statistics Theory · Mathematics 2025-12-04 Garrett G. Wen , Hong Hu , Yue M. Lu , Zhou Fan , Theodor Misiakiewicz

In the sea-land clutter classification of sky-wave over-the-horizon-radar (OTHR), the imbalanced and scarce data leads to a poor performance of the deep learning-based classification model. To solve this problem, this paper proposes an…

Systems and Control · Electrical Eng. & Systems 2023-07-19 Xiaoxuan Zhang , Zengfu Wang , Kun Lu , Quan Pan

We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We show that finding the best model under the MAPE is equivalent to doing weighted Mean Absolute Error…

Machine Learning · Statistics 2015-09-09 Arnaud De Myttenaere , Bénédicte Le Grand , Fabrice Rossi

Reliable probabilities are critical in high-risk applications, yet common calibration criteria (confidence, class-wise) are only necessary for full distributional calibration, and post-hoc methods often lack distribution-free guarantees. We…

Machine Learning · Statistics 2025-10-17 Daniil Kazantsev , Mohsen Guizani , Eric Moulines , Maxim Panov , Nikita Kotelevskii

The variational autoencoder (VAE) is a generative model with continuous latent variables where a pair of probabilistic encoder (bottom-up) and decoder (top-down) is jointly learned by stochastic gradient variational Bayes. We first…

Machine Learning · Statistics 2016-04-19 Suwon Suh , Seungjin Choi

Spatially varying coefficient (SVC) models are a type of regression model for spatial data where covariate effects vary over space. If there are several covariates, a natural question is which covariates have a spatially varying effect and…

Methodology · Statistics 2021-02-12 Jakob A. Dambon , Fabio Sigrist , Reinhard Furrer

As artificial intelligence (AI) systems are increasingly used in ethically sensitive domains such as education, healthcare, and transportation, balancing accuracy and interpretability has become a central concern. Coarse ethics (CE)…

Artificial Intelligence · Computer Science 2026-03-10 Takashi Izumo

Multimodal Retrieval-Augmented Generation (Visual RAG) significantly advances question answering by integrating visual and textual evidence. Yet, current evaluations fail to systematically account for query difficulty and ambiguity. We…

Computation and Language · Computer Science 2026-01-14 Yuelyu Ji , Wuwei Lan , Patrick NG

A general-purpose computational homogenization framework is proposed for the nonlinear dynamic analysis of membranes exhibiting complex microscale and/or mesoscale heterogeneity characterized by in-plane periodicity that cannot be…

Computational Engineering, Finance, and Science · Computer Science 2021-01-28 Philip Avery , Daniel Z. Huang , Wanli He , Johanna Ehlers , Armen Derkevorkian , Charbel Farhat

Time-series analysis is often affected by missing data, a common problem across several fields, including healthcare and environmental monitoring. Multiple Imputation by Chained Equations (MICE) has been prominent for imputing missing…

Machine Learning · Statistics 2026-04-10 Amuche Ibenegbu , Pierre Lafaye de Micheaux , Rohitash Chandra