English
Related papers

Related papers: Kendall's tau estimator for bivariate zero-inflate…

200 papers

This paper addresses the challenging computational problem of estimating intractable expectations over discrete domains. Existing approaches, including Monte Carlo and Russian Roulette estimators, are consistent but often require a large…

Machine Learning · Statistics 2025-12-19 Sophia Seulkee Kang , François-Xavier Briol , Toni Karvonen , Zonghao Chen

On the basis of Nelson-Aalen nonparametric estimator of the cumulative distribution function, we provide a weak approximation to tail product-limit process for randomly right-censored heavy-tailed data. In this context, a new consistent…

Statistics Theory · Mathematics 2016-07-25 Brahim Brahimi , Djamel Meraghni , Abdelhakim Necir

In this paper, we establish an exponential inequality for U-statistics of i.i.d. data, varying kernel and taking values in a separable Hilbert space. The bound are expressed as a sum of an exponential term plus an other one involving the…

Probability · Mathematics 2024-09-19 Davide Giraudo

Determining the strength of non-linear statistical dependencies between two variables is a crucial matter in many research fields. The established measure for quantifying such relations is the mutual information. However, estimating mutual…

Data Analysis, Statistics and Probability · Physics 2019-07-24 Damián G. Hernández , Inés Samengo

We study nonparametric estimation of univariate cumulative distribution functions (CDFs) pertaining to data missing at random. The proposed estimators smooth the inverse probability weighted (IPW) empirical CDF with the Bernstein operator,…

Statistics Theory · Mathematics 2026-03-30 Rihab Gharbi , Wissem Jedidi , Salah Khardani , Frédéric Ouimet

We propose a general semi-supervised inference framework focused on the estimation of the population mean. As usual in semi-supervised settings, there exists an unlabeled sample of covariate vectors and a labeled sample consisting of…

Methodology · Statistics 2018-08-15 Anru Zhang , Lawrence D. Brown , T. Tony Cai

An efficient monotone data augmentation (MDA) algorithm is proposed for missing data imputation for incomplete multivariate nonnormal data that may contain variables of different types, and are modeled by a sequence of regression models…

Methodology · Statistics 2018-11-21 Yongqiang Tang

Uncertainty quantification is essential in decision-making, especially when joint distributions of random variables are involved. While conformal prediction provides distribution-free prediction sets with valid coverage guarantees, it…

Machine Learning · Computer Science 2025-01-03 Rui Luo , Zhixin Zhou

Count-compositional data arise in many different fields, including high-throughput sequencing experiments, ecological surveys, and palaeoclimate studies, where a common, important goal is to understand how covariates relate to the observed…

Methodology · Statistics 2026-04-10 André F. B. Menezes , Andrew C. Parnell , Keefe Murphy

Recommender systems often suffer from selection bias as users tend to rate their preferred items. The datasets collected under such conditions exhibit entries missing not at random and thus are not randomized-controlled trials representing…

Information Retrieval · Computer Science 2024-03-05 Wonbin Kweon , Hwanjo Yu

Hiary [3] has presented an algorithm which allows to evaluate the truncated theta function $\sum_{k=0}^n \exp(2\pi \i (zk+\tau k^2))$ to within $\pm \epsilon$ in $O(\ln(\tfrac{n}{\epsilon})^{\kappa})$ arithmetic operations for any real $z$…

Number Theory · Mathematics 2014-03-25 Alexey Kuznetsov

We present a new method for estimating the frontier of a sample. The estimator is based on a local polynomial regression on the power-transformed data. We assume that the exponent of the transformation goes to infinity while the bandwidth…

Methodology · Statistics 2011-04-04 Stéphane Girard , Pierre Jacob

Worst-case bounds on the expected shortfall risk given only limited information on the distribution of the random variables has been studied extensively in the literature. In this paper, we develop a new worst-case bound on the expected…

Risk Management · Quantitative Finance 2017-02-12 Anulekha Dhara , Bikramjit Das , Karthik Natarajan

Randomized controlled trials generate experimental variation that can credibly identify causal effects, but often suffer from limited scale, while observational datasets are large, but often violate desired identification assumptions. To…

Econometrics · Economics 2023-12-27 George Z. Gui

The Negative Binomial distribution becomes highly skewed under extreme dispersion. Even at moderately large sample sizes, the sample mean exhibits a heavy right tail. The standard Normal approximation often does not provide adequate…

Methodology · Statistics 2015-03-13 David Shilane , Derek Bean

Niederreiter [H.Niederreiter, Error bounds for quasi-Monte Carlo integration with uniform point sets, Journal of computational and applied mathematics 150 (2003), 283-292] established new bounds for quasi-Monte Carlo integration for nodes…

Number Theory · Mathematics 2010-12-01 Su Hu , Yan Li

We study tail estimation in Pareto-like settings for datasets with a high percentage of randomly right-censored data, and where some expert information on the tail index is available for the censored observations. This setting arises for…

Applications · Statistics 2019-11-13 Martin Bladt , Hansjoerg Albrecher , Jan Beirlant

The Tweedie generalized linear models are commonly applied in the insurance industry to analyze semicontinuous claim data. For better prediction of the aggregated claim size, the mean and dispersion of the Tweedie model are often estimated…

Methodology · Statistics 2024-05-27 Yuwen Gu

In this paper, we develop a general theory of truncated inverse binomial sampling. In this theory, the fixed-size sampling and inverse binomial sampling are accommodated as special cases. In particular, the classical Chernoff-Hoeffding…

Statistics Theory · Mathematics 2019-08-20 Xinjia Chen

Surveys usually suffer from non-response, which decreases the effective sample size. Item non-response is typically handled by means of some form of random imputation if we wish to preserve the distribution of the imputed variable. This…

Methodology · Statistics 2017-08-04 Guillaume Chauvet , Wilfried Do Paco