English
Related papers

Related papers: A Simpson Based Estimation Approach for the Overla…

200 papers

This paper addresses the problem of estimating the containment and similarity between two sets using only random samples from each set, without relying on sketches of full sets. The study introduces a binomial model for predicting the…

Computation · Statistics 2025-07-22 Pranav Joshi

Faced with massive data, subsampling is a commonly used technique to improve computational efficiency, and using nonuniform subsampling probabilities is an effective approach to improve estimation efficiency. For computational efficiency,…

Statistics Theory · Mathematics 2022-05-19 Jing Wang , Jiahui Zou , HaiYing Wang

The Simpson's formula is obtained by approximating the integral of a function on some interval by the integral of the quadratic polynomial determined by the function. However, a multidimensional analogue of the formula has not been given as…

Mathematical Physics · Physics 2015-05-14 Kazuyuki Fujii

We give the first rigorous proof of the convergence of Riemannian Hamiltonian Monte Carlo, a general (and practical) method for sampling Gibbs distributions. Our analysis shows that the rate of convergence is bounded in terms of natural…

Data Structures and Algorithms · Computer Science 2017-10-18 Yin Tat Lee , Santosh S. Vempala

For massive data, the family of subsampling algorithms is popular to downsize the data volume and reduce computational burden. Existing studies focus on approximating the ordinary least squares estimate in linear regression, where…

Computation · Statistics 2019-06-27 HaiYing Wang , Rong Zhu , Ping Ma

The k-nearest-neighbour procedure is a well-known deterministic method used in supervised classification. This paper proposes a reassessment of this approach as a statistical technique derived from a proper probabilistic model; in…

Computation · Statistics 2008-02-12 Lionel Cucala , Jean-Michel Marin , Christian Robert , Mike Titterington

Community detection is a fundamental problem in network analysis which is made more challenging by overlaps between communities which often occur in practice. Here we propose a general, flexible, and interpretable generative model for…

Machine Learning · Statistics 2015-03-16 Yuan Zhang , Elizaveta Levina , Ji Zhu

We prove upper and lower bounds for the threshold of the q-overlap-k-Exact cover problem. These results are motivated by the one-step replica symmetry breaking approach of Statistical Physics, and the hope of using an approach based on that…

Discrete Mathematics · Computer Science 2023-09-26 Gabriel Istrate , Romeo Negrea

Nonprobability (convenience) samples are increasingly sought to stabilize estimations for one or more population variables of interest that are performed using a randomized survey (reference) sample by increasing the effective sample size.…

The damage spreading method (DS) provided a useful tool to obtain analytical results of the thermodynamics and stability of the 2D Ising model --amongst many others--, but it suffered both from ambiguities in its results and from large…

Statistical Mechanics · Physics 2009-11-13 A. Ferrera , B. Luque , L. Lacasa , E. Valero

Nonuniform subsampling methods are effective to reduce computational burden and maintain estimation efficiency for massive data. Existing methods mostly focus on subsampling with replacement due to its high computational efficiency. If the…

Methodology · Statistics 2021-07-06 Jun Yu , HaiYing Wang , Mingyao Ai , Huiming Zhang

Subsampling is an effective approach to alleviate the computational burden associated with large-scale datasets. Nevertheless, existing subsampling estimators incur a substantial loss in estimation efficiency compared to estimators based on…

Methodology · Statistics 2025-09-25 Miaomiao Su , Ruoyu Wang

This paper proposes asymptotically distribution-free inference methods for comparing a broad range of welfare indices across dependent samples, including those employed in inequality, poverty, and risk analysis. Two distinct situations are…

Econometrics · Economics 2025-12-29 Jean-Marie Dufour , Tianyu He

In stochastic combinatorial optimization, algorithms differ in their adaptivity: whether or not they query realized randomness and adapt to it. Dean et al. (FOCS '04) formalize the adaptivity gap, which compares the performance of fully…

Data Structures and Algorithms · Computer Science 2026-03-03 Zohar Barak , Inbal Talgam-Cohen

We consider the question of learning the natural parameters of a $k$ parameter minimal exponential family from i.i.d. samples in a computationally and statistically efficient manner. We focus on the setting where the support as well as the…

Machine Learning · Computer Science 2021-11-01 Abhin Shah , Devavrat Shah , Gregory W. Wornell

A new nonparametric approach, based on a decision tree algorithm, is proposed to calculate the overlap between two probability distributions. The devised framework is described analytically and numerically. The convergence of the estimated…

Statistics Theory · Mathematics 2022-11-28 Hisashi Johno , Kazunori Nakamoto

Multivariate normal mixtures provide a flexible model for high-dimensional data. They are widely used in statistical genetics, statistical finance, and other disciplines. Due to the unboundedness of the likelihood function, classical…

Statistics Theory · Mathematics 2008-05-27 Jiahua Chen , Xianming Tan

In this paper we present a new measure for the overlap of two density functions which provides motivation and interpretation currently lacking with benchmark measures based on the proportion of similar response, also known as the overlap…

Methodology · Statistics 2021-06-08 Stephen G Walker

This work presents a statistically principled method for estimating the required number of instances in the experimental comparison of multiple algorithms on a given problem class of interest. This approach generalises earlier results by…

Methodology · Statistics 2019-08-06 Felipe Campelo , Elizabeth F. Wanner

Simpson's paradox, a long-standing statistical phenomenon, describes the reversal of an observed association when data are disaggregated into sub-populations. It has critical implications across statistics, epidemiology, economics, and…

Databases · Computer Science 2025-11-04 Yi Yang , Jian Pei , Jun Yang , Jichun Xie