English
Related papers

Related papers: Higher Criticism to Compare Two Large Frequency Ta…

200 papers

Refining one's hypotheses in the light of data is a common scientific practice; however, the dependency on the data introduces selection bias and can lead to specious statistical analysis. An approach for addressing this is via conditioning…

Machine Learning · Computer Science 2020-03-03 Jen Ning Lim , Makoto Yamada , Wittawat Jitkrittum , Yoshikazu Terada , Shigeyuki Matsui , Hidetoshi Shimodaira

We consider a sequence of Hawkes processes whose excitation measures may depend on the generation, and study its scaling limits in the near-unstable limiting regime. The limiting random measures, characterized via a nonlinear convolutional…

Probability · Mathematics 2026-04-08 Tristan Pace , Gordan Zitkovic

Motivated by the increasing use of kernel-based metrics for high-dimensional and large-scale data, we study the asymptotic behavior of kernel two-sample tests when the dimension and sample sizes both diverge to infinity. We focus on the…

Statistics Theory · Mathematics 2024-10-31 Jian Yan , Xianyang Zhang

The local number variance associated with a spherical sampling window of radius $R$ enables a classification of many-particle systems in $d$-dimensional Euclidean space according to the degree to which large-scale density fluctuations are…

Statistical Mechanics · Physics 2021-05-12 Salvatore Torquato , Jaeuk Kim , Michael A. Klatt

We develop randomization-based tests for heterogeneous treatment effects in the presence of network interference. Leveraging the exposure mapping framework, we study a broad class of null hypotheses that represent various forms of constant…

Econometrics · Economics 2025-06-25 Julius Owusu

We investigate the advantage of coherent superposition of two different coded channels in quantum metrology. In a continuous variable system, we show that the Heisenberg limit $1/N$ can be beaten by the coherent superposition without the…

Quantum Physics · Physics 2021-12-15 Dong Xie , Chunling Xu , An Min Wang

In this paper, we study various models for random combinatorial partitions using large deviation analysis for diverging scale of the reference process. Scaling limits of similar models have been studied recently \cite{FSa,FSb} going back to…

Probability · Mathematics 2022-06-22 Stefan Adams , Matthew Dickson

Two semimetrics on probability distributions are proposed, given as the sum of differences of expectations of analytic functions evaluated at spatial or frequency locations (i.e, features). The features are chosen so as to maximize the…

Machine Learning · Statistics 2016-10-31 Wittawat Jitkrittum , Zoltan Szabo , Kacper Chwialkowski , Arthur Gretton

Characterizing samples that are difficult to learn from is crucial to developing highly performant ML models. This has led to numerous Hardness Characterization Methods (HCMs) that aim to identify "hard" samples. However, there is a lack of…

Machine Learning · Computer Science 2024-03-08 Nabeel Seedat , Fergus Imrie , Mihaela van der Schaar

A panel dataset satisfies marginal homogeneity if the time-specific marginal distributions are homogeneous or time-invariant. Marginal homogeneity is relevant in many economic settings, including dynamic discrete games,…

Econometrics · Economics 2025-12-08 Federico Bugni , Jackson Bunting , Muyang Ren

Large-scale network inference with uncertainty quantification has important applications in natural, social, and medical sciences. The recent work of Fan, Fan, Han and Lv (2022) introduced a general framework of statistical inference on…

Machine Learning · Statistics 2022-11-02 Jianqing Fan , Yingying Fan , Jinchi Lv , Fan Yang

From a stability perspective, a renewable generation (RG)-rich power system is a constrained system. As the quasistability boundary of a constrained system is structurally very different from that of an unconstrained system, finding the…

Systems and Control · Computer Science 2020-02-05 Chetan Mishra , Anamitra Pal , Virgilio A. Centeno

Statistical inferences for sample correlation matrices are important in high dimensional data analysis. Motivated by this, this paper establishes a new central limit theorem (CLT) for a linear spectral statistic (LSS) of high dimensional…

Statistics Theory · Mathematics 2014-11-04 Jiti Gao , Xiao Han , Guangming Pan , Yanrong Yang

Large-scale multiple testing problems require the simultaneous assessment of many p-values. This paper compares several methods to assess the evidence in multiple binomial counts of p-values: the maximum of the binomial counts after…

Methodology · Statistics 2014-02-26 Guenther Walther

Data-driven most powerful tests are statistical hypothesis decision-making tools that deliver the greatest power against a fixed null hypothesis among all corresponding data-based tests of a given size. When the underlying data…

Statistics Theory · Mathematics 2023-03-15 Albert Vexler , Alan D. Hutson

We study the problem of two-sample comparison with categorical data when the contingency table is sparsely populated. In modern applications, the number of categories is often comparable to the sample size, causing existing methods to have…

Methodology · Statistics 2014-08-14 Hao Chen , Nancy R. Zhang

Testing the equality of the covariance matrices of two high-dimensional samples is a fundamental inference problem in statistics. Several tests have been proposed but they are either too liberal or too conservative when the required…

Statistics Theory · Mathematics 2023-01-04 Jin-Ting Zhang , Jingyi Wang , Tianming Zhu

We present a general framework for hypothesis testing on distributions of sets of individual examples. Sets may represent many common data sources such as groups of observations in time series, collections of words in text or a batch of…

Methodology · Statistics 2021-02-03 Alexis Bellot , Mihaela van der Schaar

In typical high dimensional statistical inference problems, confidence intervals and hypothesis tests are performed for a low dimensional subset of model parameters under the assumption that the parameters of interest are unconstrained.…

Methodology · Statistics 2019-11-19 Ming Yu , Varun Gupta , Mladen Kolar

We propose a novel algorithm for large-scale regression problems named histogram transform ensembles (HTE), composed of random rotations, stretchings, and translations. First of all, we investigate the theoretical properties of HTE when the…

Machine Learning · Statistics 2019-12-11 Hanyuan Hang , Zhouchen Lin , Xiaoyu Liu , Hongwei Wen