English
Related papers

Related papers: Penalized Quasi-likelihood for High-dimensional Lo…

200 papers

The Ising model is a useful tool for studying complex interactions within a system. The estimation of such a model, however, is rather challenging, especially in the presence of high-dimensional parameters. In this work, we propose…

Statistics Theory · Mathematics 2012-08-20 Lingzhou Xue , Hui Zou , Tianxi Cai

Establishing a low-dimensional representation of the data leads to efficient data learning strategies. In many cases, the reduced dimension needs to be explicitly stated and estimated from the data. We explore the estimation of dimension in…

Methodology · Statistics 2022-02-10 Wei Q. Deng , Radu V. Craiu

Quantile regression is often used when a comprehensive relationship between a response variable and one or more explanatory variables is desired. The traditional frequentists' approach to quantile regression has been well developed around…

Statistics Theory · Mathematics 2015-06-03 Yang Feng , Yuguo Chen , Xuming He

We present new methods to estimate causal effects retrospectively from micro data with the assistance of a machine learning ensemble. This approach overcomes two important limitations in conventional methods like regression modeling or…

Machine Learning · Statistics 2016-07-12 Cyrus Samii , Laura Paler , Sarah Zukerman Daly

Computing averages over a target probability density by statistical re-weighting of a set of samples with a different distribution is a strategy which is commonly adopted in fields as diverse as atomistic simulation and finance. Here we…

Chemical Physics · Physics 2012-02-21 Michele Ceriotti , Guy A. R. Brain , Oliver Riordan , David E. Manolopoulos

Precisely estimating out-of-sample upper quantiles is very important in risk assessment and in engineering practice for structural design to prevent a greater disaster. For this purpose, the generalized extreme value (GEV) distribution has…

Methodology · Statistics 2025-12-24 Yonggwan Shin , Yire Shin , Jihong Park , Jeong-Soo Park

Many experiments in medicine and ecology can be conveniently modeled by finite Gaussian mixtures but face the problem of dealing with small data sets. We propose a robust version of the estimator based on self-regression and sparsity…

Computation · Statistics 2014-10-09 Stephane Chretien

Gaussian graphical modeling has been widely used to explore various network structures, such as gene regulatory networks and social networks. We often use a penalized maximum likelihood approach with the $L_1$ penalty for learning a…

Methodology · Statistics 2017-06-13 Kei Hirose , Hironori Fujisawa , Jun Sese

Auto-regressive sequence generative models trained by Maximum Likelihood Estimation suffer the exposure bias problem in practical finite sample scenarios. The crux is that the number of training samples for Maximum Likelihood Estimation is…

Machine Learning · Statistics 2020-07-14 Yuxuan Song , Ning Miao , Hao Zhou , Lantao Yu , Mingxuan Wang , Lei Li

Introduction In analysis of time-to-event outcomes, a mixture cure (MC) model is preferred over a standard survival model when the sample includes individuals who will never experience the event of interest. Motivated by a cohort study of…

Applications · Statistics 2026-04-02 Changchang Xu , Laurent Briollais , Irene L Andrulis , Shelley B Bull

We propose a doubly robust estimator for the average treatment effect in high dimensional low sample size observational studies, where contamination and model misspecification pose serious inferential challenges. The estimator combines…

Methodology · Statistics 2025-11-04 Byeonghee Lee , Sangwook Kang , Ju-Hyun Park , Saebom Jeon , Joonsung Kang

Clustering is a fundamental tool in statistical machine learning in the presence of heterogeneous data. Most recent results focus primarily on optimal mislabeling guarantees when data are distributed around centroids with sub-Gaussian…

Statistics Theory · Mathematics 2024-10-24 Soham Jana , Jianqing Fan , Sanjeev Kulkarni

Mixtures-of-Experts models and their maximum likelihood estimation (MLE) via the EM algorithm have been thoroughly studied in the statistics and machine learning literature. They are subject of a growing investigation in the context of…

Machine Learning · Statistics 2019-09-13 Faïcel Chamroukhi , Florian Lecocq , Hien D. Nguyen

This article concerns the dimension reduction in regression for large data set. We introduce a new method based on the sliced inverse regression approach, called cluster-based regularized sliced inverse regression. Our method not only keeps…

Applications · Statistics 2013-12-03 Yue Yu , Zhihong Chen , Jie Yang

An approximate method for conducting resampling in Lasso, the $\ell_1$ penalized linear regression, in a semi-analytic manner is developed, whereby the average over the resampled datasets is directly computed without repeated numerical…

Machine Learning · Statistics 2018-12-11 Tomoyuki Obuchi , Yoshiyuki Kabashima

We consider joint estimation of multiple graphical models arising from heterogeneous and high-dimensional observations. Unlike most previous approaches which assume that the cluster structure is given in advance, an appealing feature of our…

Machine Learning · Statistics 2018-01-16 Botao Hao , Will Wei Sun , Yufeng Liu , Guang Cheng

Variable clustering is important for explanatory analysis. However, only few dedicated methods for variable clustering with the Gaussian graphical model have been proposed. Even more severe, small insignificant partial correlations due to…

Applications · Statistics 2018-06-18 Daniel Andrade , Akiko Takeda , Kenji Fukumizu

We devise survey-weighted pseudo posterior distribution estimators under two-stage informative sampling of both primary clusters and secondary nested units for a one-way analysis of variance (ANOVA) population generating model as a simple…

Methodology · Statistics 2023-05-16 Terrance D. Savitsky , Matthew R. Williams , Sanvesh Srivastava

High covariate dimensionality is increasingly occurrent in model estimation, and existing techniques to address this issue typically require sparsity or discrete heterogeneity of the \emph{unobservable} parameter vector. However, neither…

Econometrics · Economics 2025-07-31 Abdul-Nasah Soale , Emmanuel Selorm Tsyawo

Longitudinal data are commonly encountered in biomedical research, including randomized trials and retrospective cohort studies. Subjects are typically followed over a period of time and may be scheduled for follow-up at pre-determined time…

Methodology · Statistics 2025-10-23 George Stefan , Eleanor Pullenayegum
‹ Prev 1 4 5 6 7 8 10 Next ›