English
Related papers

Related papers: Asymptotically Optimal Sequential Confidence Inter…

200 papers

We consider nonparametric testing in a non-asymptotic framework. Our statistical guarantees are exact in the sense that Type I and II errors are controlled for any finite sample size. Meanwhile, one proposed test is shown to achieve minimax…

Statistics Theory · Mathematics 2017-02-07 Yun Yang , Zuofeng Shang , Guang Cheng

Standard empirical risk minimization (ERM) models may prioritize learning spurious correlations between spurious features and true labels, leading to poor accuracy on groups where these correlations do not hold. Mitigating this issue often…

Machine Learning · Computer Science 2024-06-05 Yujin Han , Difan Zou

Estimating the number of clusters (K) is a critical and often difficult task in cluster analysis. Many methods have been proposed to estimate K, including some top performers using resampling approach. When performing cluster analysis in…

Methodology · Statistics 2019-09-05 Yujia Li , Xiangrui Zeng , Chien-Wei Lin , George Tseng

A balanced sampling design should always be the adopted strategies if auxiliary information is available. Besides, integrating a stratified structure of the population in the sampling process can considerably reduce the variance of the…

Methodology · Statistics 2022-06-03 Raphaël Jauslin , Esther Eustache , Yves Tillé

Uniform sampling and approximate counting are fundamental primitives for modern database applications, ranging from query optimization to approximate query processing. While recent breakthroughs have established optimal sampling and…

Databases · Computer Science 2026-05-13 Xiao Hu , Jinchao Huang

In the analysis of cluster data, the regression coefficients are frequently assumed to be the same across all clusters. This hampers the ability to study the varying impacts of factors on each cluster. In this paper, a semiparametric model…

Statistics Theory · Mathematics 2009-08-25 Wenyang Zhang , Jianqing Fan , Yan Sun

We consider quantile estimation in a semi-supervised setting, characterized by two available data sets: (i) a small or moderate sized labeled data set containing observations for a response and a set of possibly high dimensional covariates,…

Methodology · Statistics 2024-08-15 Abhishek Chakrabortty , Guorong Dai , Raymond J. Carroll

We study the high-dimensional linear regression problem with categorical predictors that have many levels. We propose a new estimation approach, which performs model compression via two mechanisms by simultaneously encouraging (a)…

Methodology · Statistics 2026-03-30 Kayhan Behdin , Riade Benbaki , Peter Radchenko , Rahul Mazumder

We consider the problem of simultaneous variable selection and estimation in additive, partially linear models for longitudinal/clustered data. We propose an estimation procedure via polynomial splines to estimate the nonparametric…

Statistics Theory · Mathematics 2013-02-04 Shujie Ma , Qiongxia Song , Li Wang

This paper focuses on the problem of recursive nonlinear least squares parameter estimation in multi-agent networks, in which the individual agents observe sequentially over time an independent and identically distributed (i.i.d.)…

Optimization and Control · Mathematics 2016-10-20 Anit Kumar Sahu , Soummya Kar , Jose' M. F. Moura , H. Vincent Poor

In this paper we apply a two-stage sequential design to item calibration problems under a three-parameter logistic model assumption. The measurement errors of the estimates of the latent trait levels of examinees are considered in our…

Applications · Statistics 2013-05-23 Yuan-chin Ivan Chang

Internal measures that are used to assess the quality of a clustering usually take into account intra-group and/or inter-group criteria. There are many papers in the literature that propose algorithms with provable approximation guarantees…

Machine Learning · Computer Science 2024-01-17 Eduardo S. Laber , Lucas Murtinho

When dealing with very large datasets of functional data, survey sampling approaches are useful in order to obtain estimators of simple functional quantities, without being obliged to store all the data. We propose here a Horvitz--Thompson…

Methodology · Statistics 2011-11-29 Hervé Cardot , Etienne Josserand

Neural network-based methods for (un)conditional density estimation have recently gained substantial attention, as various neural density estimators have outperformed classical approaches in real-data experiments. Despite these empirical…

Machine Learning · Statistics 2025-10-02 Dehao Dai , Jianqing Fan , Yihong Gu , Debarghya Mukherjee

This Ph.D. thesis explores approximations and regularity for the Heston stochastic volatility model through three interconnected works. The first work focuses on developing high-order weak approximations for the Cox-Ingersoll-Ross (CIR)…

Numerical Analysis · Mathematics 2025-05-01 Edoardo Lombardo

The Bayesian formulation of sequentially testing $M \ge 3$ hypotheses is studied in the context of a decentralized sensor network system. In such a system, local sensors observe raw observations and send quantized sensor messages to a…

Statistics Theory · Mathematics 2016-11-15 Yan Wang , Yajun Mei

We present the first method for assessing the relevance of a model-based clustering result in a general framework. Standard validation criteria, like the adjusted Rand index, rely on external labels to assess partition accuracy;…

Statistics Theory · Mathematics 2026-03-30 Salima El Kolei , Matthieu Marbac

Tail Gini functional is a measure of tail risk variability for systemic risks, and has many applications in banking, finance and insurance. Meanwhile, there is growing attention on aymptotic independent pairs in quantitative risk…

Methodology · Statistics 2023-09-13 Zhaowen Wang , Liujun Chen , Deyuan Li

Categorical Gini Correlation (CGC), introduced by Dang et al. (2020), is a novel dependence measure designed to quantify the association between a numerical variable and a categorical variable. It has appealing properties compared to…

Methodology · Statistics 2026-05-12 Sameera Hewage

This study introduces a general semiparametric clusterwise index distribution model to analyze how latent clusters affect the covariate-response relationships. By employing sufficient dimension reduction to account for the effects of…

Methodology · Statistics 2025-09-30 Jen-Chieh Teng , Chin-Tsang Chiang