English
Related papers

Related papers: Better lower bounds for missing species: improved …

200 papers

Existing identification and estimation methods for semiparametric sample selection models rely heavily on exclusion restrictions. However, it is difficult in practice to find a credible excluded variable that has a correlation with…

Econometrics · Economics 2024-12-03 Zhewen Pan , Yifan Zhang

Generalized linear models are flexible tools for the analysis of diverse datasets, but the classical formulation requires that the parametric component is correctly specified and the data contain no atypical observations. To address these…

Methodology · Statistics 2023-04-21 Ioannis Kalogridis , Gerda Claeskens , Stefan Van Aelst

The maximum likelihood estimation is computationally demanding for large datasets, particularly when the likelihood function includes integrals. Subsampling can reduce the computational burden, but it often results in efficiency loss.This…

Methodology · Statistics 2026-04-27 Miaomiao Su , Qihua Wang , Ruoyu Wang

Complex biological processes are usually experimented along time among a collection of individuals. Longitudinal data are then available and the statistical challenge is to better understand the underlying biological mechanisms. The…

Statistics Theory · Mathematics 2015-06-11 Pierre Barbillon , Célia Barthélémy , Adeline Samson

In the world of multivariate extremes, estimation of the dependence structure still presents a challenge and an interesting problem. A procedure for the bivariate case is presented that opens the road to a similar way of handling the…

Statistics Theory · Mathematics 2008-11-14 John H. J. Einmahl , Andrea Krajina , Johan Segers

Multi-category data arise in diverse fields including marketing, chemistry, public policy, genomics, political science, and ecology. We consider the problem of estimating ratios of category-specific means in a fully nonparametric setting,…

Methodology · Statistics 2025-10-29 Grant Hopkins , Sarah Teichman , Ellen Graham , Amy D Willis

Probability density estimation is a core problem of statistics and signal processing. Moment methods are an important means of density estimation, but they are generally strongly dependent on the choice of feasible functions, which severely…

Machine Learning · Statistics 2023-07-06 Guangyu Wu , Anders Lindquist

For many tasks of data analysis, we may only have the information of the explanatory variable and the evaluation of the response values are quite expensive. While it is impractical or too costly to obtain the responses of all units, a…

Computation · Statistics 2023-04-07 Wei Zheng , Ting Tian , Xueqin Wang

The paper describes a new class of capture-recapture models for closed populations when individual covariates are available. The novelty consists in combining a latent class model for the distribution of the capture history, where the class…

Methodology · Statistics 2021-11-08 Antonio Forcina , Francesco Bartolucci

We developed a statistical inference method applicable to a broad range of generalized linear models (GLMs) in high-dimensional settings, where the number of unknown coefficients scales proportionally with the sample size. Although a…

Statistics Theory · Mathematics 2024-05-24 Kazuma Sawaya , Yoshimasa Uematsu , Masaaki Imaizumi

This paper proposes a linear categorical random coefficient model, in which the random coefficients follow parametric categorical distributions. The distributional parameters are identified based on a linear recurrence structure of moments…

Econometrics · Economics 2023-03-01 Zhan Gao , M. Hashem Pesaran

The paper aims at reconsidering the famous Le Cam LAN theory. The main features of the approach which make it different from the classical one are as follows: (1) the study is nonasymptotic, that is, the sample size is fixed and does not…

Statistics Theory · Mathematics 2013-03-06 Vladimir Spokoiny

In model checking for regressions, nonparametric estimation-based tests usually have tractable limiting null distributions and are sensitive to oscillating alternative models, but suffer from the curse of dimensionality. In contrast,…

Methodology · Statistics 2019-03-12 Lingzhu Li , Xuehu Zhu , Lixing Zhu

In this paper, a shrinkage estimator for the population mean is proposed under known quadratic loss functions with unknown covariance matrices. The new estimator is non-parametric in the sense that it does not assume a specific parametric…

Methodology · Statistics 2014-11-07 Cheng Wang , Tiejun Tong , Longbing Cao , Baiqi Miao

The Cox proportional hazards model is widely used in survival analysis to model time-to-event data. However, it faces significant computational challenges in the era of large-scale data, particularly when dealing with time-dependent…

Methodology · Statistics 2025-01-14 Miaomiao Su , Ruoyu Wang

Hierarchical models allow for heterogeneous behaviours in a population while simultaneously borrowing estimation strength across all subpopulations. Unfortunately, existing likelihood-based methods for fitting hierarchical models have high…

Methodology · Statistics 2015-12-16 Patrick O. Perry

Change-point detection has been a classical problem in statistics and econometrics. This work focuses on the problem of detecting abrupt distributional changes in the data-generating distribution of a sequence of high-dimensional…

Methodology · Statistics 2021-05-20 Shubhadeep Chakraborty , Xianyang Zhang

Predicting the timing and occurrence of events is a major focus of data science applications, especially in the context of biomedical research. Performance for models estimating these outcomes, often referred to as time-to-event or survival…

Methodology · Statistics 2024-06-07 Ying Jin , Andrew Leroux

This paper considers nonparametric identification and estimation of the regression function when a covariate is mismeasured. The measurement error need not be classical. Employing the small measurement error approximation, we establish…

Econometrics · Economics 2024-03-19 Kirill S. Evdokimov , Andrei Zeleneev

Mixture models are a fundamental tool in applied statistics and machine learning for treating data taken from multiple subpopulations. The current practice for estimating the parameters of such models relies on local search heuristics…

Machine Learning · Computer Science 2012-09-07 Animashree Anandkumar , Daniel Hsu , Sham M. Kakade