English
Related papers

Related papers: Pseudo-$R^2$ statistics under complex sampling

200 papers

We consider the problem of parametric statistical inference when likelihood computations are prohibitively expensive but sampling from the model is possible. Several so-called likelihood-free methods have been developed to perform inference…

Machine Learning · Statistics 2020-09-14 Owen Thomas , Ritabrata Dutta , Jukka Corander , Samuel Kaski , Michael U. Gutmann

Robust statistics traditionally focuses on outliers, or perturbations in total variation distance. However, a dataset could be corrupted in many other ways, such as systematic measurement errors and missing covariates. We generalize the…

Statistics Theory · Mathematics 2020-12-15 Banghua Zhu , Jiantao Jiao , Jacob Steinhardt

We consider the variable selection problem of generalized linear models (GLMs). Stability selection (SS) is a promising method proposed for solving this problem. Although SS provides practical variable selection criteria, it is…

Machine Learning · Statistics 2025-08-06 Takashi Takahashi , Yoshiyuki Kabashima

The aim of this note is to state a couple of general results about the properties of the penalized maximum likelihood estimators (pMLE) and of the posterior distribution for parametric models in a non-asymptotic setup and for possibly large…

Statistics Theory · Mathematics 2022-12-13 Vladimir Spokoiny

This paper studies Cox's regression hazard model with an unobservable random frailty where no specific distribution is postulated for the frailty variable, and the marginal lifetime distribution allows both parametric and non-parametric…

Methodology · Statistics 2015-10-09 Vahed Maroufy , Paul Marriott

A popular technique for selecting and tuning machine learning estimators is cross-validation. Cross-validation evaluates overall model fit, usually in terms of predictive accuracy. In causal inference, the optimal choice of estimator…

Methodology · Statistics 2021-07-07 Dominik Rothenhäusler

The goal of subsampling is to select an informative subset of all observations, when using the full data for statistical analysis is not viable. We construct locally $ D $-optimal subsampling designs under a Poisson regression model with a…

Statistics Theory · Mathematics 2024-03-28 Torsten Reuter , Rainer Schwabe

The standard linear and logistic regression models assume that the response variables are independent, but share the same linear relationship to their corresponding vectors of covariates. The assumption that the response variables are…

Machine Learning · Computer Science 2019-10-09 Constantinos Daskalakis , Nishanth Dikkala , Ioannis Panageas

We propose elliptical graphical models based on conditional uncorrelatedness as a general- ization of Gaussian graphical models by letting the population distribution be elliptical instead of normal, allowing the fitting of data with…

Methodology · Statistics 2015-06-16 Daniel Vogel , Roland Fried

At least two, different approaches to define and solve statistical models for the analysis of economic systems exist: the typical, econometric one, interpreting the Gravity Model specification as the expected link weight of an arbitrary…

Physics and Society · Physics 2023-11-06 Marzio Di Vece , Diego Garlaschelli , Tiziano Squartini

Different statistical samples (e.g., from different locations) offer populations and learning systems observations with distinct statistical properties. Samples under (1) 'Unconfounded' growth preserve systems' ability to determine the…

Methodology · Statistics 2025-08-05 Andre F. Ribeiro

This paper considers the problem of estimation in the generalized semiparametric model for longitudinal data when the number of parameters diverges with the sample size. A penalization type of generalized estimating equation method is…

Methodology · Statistics 2020-06-09 M. Taavoni , M. Arashi

The problem of efficiently generating random samples from high-dimensional and non-log-concave posterior measures arising from nonlinear regression problems is considered. Extending investigations from arXiv:2009.05298, local and global…

Statistics Theory · Mathematics 2023-04-18 Jan Bohr , Richard Nickl

This work performs a non-asymptotic analysis of the generalized Lasso under the assumption of sub-exponential data. Our main results continue recent research on the benchmark case of (sub-)Gaussian sample distributions and thereby explore…

Statistics Theory · Mathematics 2023-01-18 Martin Genzel , Christian Kipp

A class of estimating functions is introduced for the regression parameter of the Cox proportional hazards model to allow unknown failure statuses on some study subjects. The consistency and asymptotic normality of the resulting estimators…

Statistics Theory · Mathematics 2007-08-22 Irene Gijbels , Danyu Lin , Zhiliang Ying

Case-cohort design, an outcome-dependent sampling design for censored survival data, is increasingly used in biomedical research. The development of asymptotic theory for a case-cohort design in the current literature primarily relies on…

Statistics Theory · Mathematics 2012-04-13 Bin Nan , Jon A. Wellner

Complicated generative models often result in a situation where computing the likelihood of observed data is intractable, while simulating from the conditional density given a parameter value is relatively easy. Approximate Bayesian…

Machine Learning · Statistics 2015-12-29 Mijung Park , Wittawat Jitkrittum , Dino Sejdinovic

We address the component-based regularisation of a multivariate Generalized Linear Mixed Model (GLMM). A set of random responses Y is modelled by a GLMM, using a set X of explanatory variables, a set T of additional covariates, and random…

Methodology · Statistics 2019-08-22 Jocelyn Chauvet , Catherine Trottier , Xavier Bry , Frederic Mortier

We consider model selection in generalized linear models (GLM) for high-dimensional data and propose a wide class of model selection criteria based on penalized maximum likelihood with a complexity penalty on the model size. We derive a…

Statistics Theory · Mathematics 2016-03-31 Felix Abramovich , Vadim Grinshtein

Recent work has focused on the problem of conducting linear regression when the number of covariates is very large, potentially greater than the sample size. To facilitate this, one useful tool is to assume that the model can be well…

Methodology · Statistics 2011-11-21 Zhou Fang