English
Related papers

Related papers: Flexible variable selection in the presence of mis…

200 papers

Generalized linear models are flexible tools for the analysis of diverse datasets, but the classical formulation requires that the parametric component is correctly specified and the data contain no atypical observations. To address these…

Methodology · Statistics 2023-04-21 Ioannis Kalogridis , Gerda Claeskens , Stefan Van Aelst

This research deals with the estimation and imputation of missing data in longitudinal models with a Poisson response variable inflated with zeros. A methodology is proposed that is based on the use of maximum likelihood, assuming that data…

Methodology · Statistics 2024-09-18 D. S. Martinez-Lobo , O. O. Melo , N. A. Cruz

While advances continue to be made in model-based clustering, challenges persist in modeling various data types such as panel data. Multivariate panel data present difficulties for clustering algorithms because they are often plagued by…

Methodology · Statistics 2024-08-26 Mackenzie R. Neal , Alexa A. Sochaniwsky , Paul D. McNicholas

This paper tackles the problem of constructing a non-parametric predictor when the latent variables are given with incomplete information. The convenient predictor for this task is the random forest algorithm in conjunction to the so-called…

Statistics Theory · Mathematics 2023-09-01 Irving Gómez-Méndez , Emilien Joly

The problem of choosing appropriate values for missing data is often encountered in the data science. We describe a novel method containing both traditional mathematics and machine learning elements for prediction (imputation) of missing…

Machine Learning · Computer Science 2025-10-13 Peteris Daugulis , Vija Vagale , Emiliano Mancini , Filippo Castiglione

Modal regression has emerged as a flexible alternative to classical regression models when the conditional mean or median are unable to adequately capture the underlying relation between a response and a predictor variable. This approach is…

Methodology · Statistics 2025-04-08 Ana Pérez-González , Tomás R. Cotos-Yáñez , Rosa M. Crujeiras

We consider a Gaussian sequence space model $X_{\lambda}=f_{\lambda} + \xi_{\lambda},$ where $\xi $ has a diagonal covariance matrix $\Sigma=\diag(\sigma_\lambda ^2)$. We consider the situation where the parameter vector $(f_{\lambda})$ is…

Statistics Theory · Mathematics 2013-12-23 Laurent Cavalier , Markus Reiß

A desirable data selection algorithm can efficiently choose the most informative samples to maximize the utility of limited annotation budgets. However, current approaches, represented by active learning methods, typically follow a…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Yichen Xie , Mingyu Ding , Masayoshi Tomizuka , Wei Zhan

Missing data is an important challenge when dealing with high dimensional data arranged in the form of an array. In this paper, we propose methods for estimation of the parameters of array variate normal probability model from partially…

Methodology · Statistics 2015-01-06 Deniz Akdemir

The importance of variable selection for clustering has been recognized for some time, and mixture models are well-established as a statistical approach to clustering. Yet, the literature on variable selection in model-based clustering…

Methodology · Statistics 2024-02-13 Mackenzie R. Neal , Paul D. McNicholas

We consider identification, inference and validation of linear panel data models when both factors and factor loadings are accounted for by a nonparametric function. This general specification encompasses rather popular models such as the…

Econometrics · Economics 2025-06-13 Juan M. Rodriguez-Poo , Alexandra Soberon , Stefan Sperlich

Semi-competing risks data arise when both non-terminal and terminal events are considered in a model. Such data with multiple events of interest are frequently encountered in medical research and clinical trials. In this framework, terminal…

Methodology · Statistics 2022-11-21 Fatemeh Mahmoudi , Xuewen Lu

Mixed-effect models are very popular for analyzing data with a hierarchical structure, e.g. repeated observations within subjects in a longitudinal design, patients nested within centers in a multicenter design. However, recently, due to…

Methodology · Statistics 2019-05-09 Abhik Ghosh , Magne Thoresen

In genetic studies, not only can the number of predictors obtained from microarray measurements be extremely large, there can also be multiple response variables. Motivated by such a situation, we consider semiparametric dimension reduction…

Methodology · Statistics 2013-09-25 Heng Lian , Shujie Ma

We investigate methods for penalized regression in the presence of missing observations. This paper introduces a method for estimating the parameters which compensates for the missing observations. We first, derive an unbiased estimator of…

Applications · Statistics 2013-10-09 Yunjin Choi , Robert Tibshirani

Statistical inference in high dimensional settings has recently attracted enormous attention within the literature. However, most published work focuses on the parametric linear regression problem. This paper considers an important…

Methodology · Statistics 2019-11-14 Qi Gao , Randy C. S. Lai , Thomas C. M. Lee , Yao Li

Missing data arises when certain values are not recorded or observed for variables of interest. However, most of the statistical theory assume complete data availability. To address incomplete databases, one approach is to fill the gaps…

This paper proposes an imputation procedure that uses the factors estimated from a tall block along with the re-rotated loadings estimated from a wide block to impute missing values in a panel of data. Assuming that a strong factor…

Econometrics · Economics 2021-08-13 Jushan Bai , Serena Ng

Many computer vision and medical imaging problems are faced with learning from large-scale datasets, with millions of observations and features. In this paper we propose a novel efficient learning scheme that tightens a sparsity constraint…

Machine Learning · Statistics 2017-02-07 Adrian Barbu , Yiyuan She , Liangjing Ding , Gary Gramajo

The main contribution of this paper is the development of a new decision tree algorithm. The proposed approach allows users to guide the algorithm through the data partitioning process. We believe this feature has many applications but in…

Machine Learning · Statistics 2020-10-27 Cédric Beaulac , Jeffrey S. Rosenthal