English
Related papers

Related papers: Two-sample test of sparse stochastic block models

200 papers

We consider a two-sample hypothesis testing problem, where the distributions are defined on the space of undirected graphs, and one has access to only one observation from each model. A motivating example for this problem is comparing the…

Estimating the asymmetric numbers of communities in multi-layer directed networks is a challenging problem due to the multi-layer structures and inherent directional asymmetry, leading to possibly different numbers of sender and receiver…

Statistics Theory · Mathematics 2026-02-26 Huan Qing

How can one determine whether a community-level treatment, such as the introduction of a social program or trade shock, alters agents' incentives to form links in a network? This paper proposes analogues of a two-sample Kolmogorov-Smirnov…

Econometrics · Economics 2020-11-24 Eric Auerbach

The log-normal distribution is one of the most common distributions used for modeling skewed and positive data. It frequently arises in many disciplines of science, specially in the biological and medical sciences. The statistical analysis…

Methodology · Statistics 2020-01-01 Ayanendranath Basu , Abhijit Mandal , Nirian Martin , Leandro Pardo

In the standard stochastic block model for networks, the probability of a connection between two nodes, often referred to as the edge probability, depends on the unobserved communities each of these nodes belongs to. We consider a flexible…

Econometrics · Economics 2024-02-27 Yuichi Kitamura , Louise Laage

To characterize the community structure in network data, researchers have introduced various block-type models, including the stochastic block model, degree-corrected stochastic block model, mixed membership block model, degree-corrected…

Methodology · Statistics 2024-09-10 Yujia Wu , Jingfei Zhang , Wei Lan , Chih-Ling Tsai

The stochastic block model is one of the most studied network models for community detection. It is well-known that most algorithms proposed for fitting the stochastic block model likelihood function cannot scale to large-scale networks.…

Methodology · Statistics 2021-08-31 Jiangzhou Wang , Jingfei Zhang , Binghui Liu , Ji Zhu , Jianhua Guo

We consider a non-projective class of inhomogeneous random graph models with interpretable parameters and a number of interesting asymptotic properties. Using the results of Bollob\'as et al. [2007], we show that i) the class of models is…

Machine Learning · Statistics 2018-10-04 Juho Lee , Lancelot F. James , Seungjin Choi , François Caron

In this paper, we address the problem of two-sample testing in the presence of missing data under a variety of missingness mechanisms. Our focus is on the well-known energy distance-based two-sample test. In addition to the standard…

Methodology · Statistics 2025-08-18 Danijel G. Aleksić , Bojana Milošević

The stochastic block model is widely used for detecting community structures in network data. How to test the goodness-of-fit of the model is one of the fundamental problems and has gained growing interests in recent years. In this article,…

Methodology · Statistics 2019-08-27 Jianwei Hu , Jingfei Zhang , Hong Qin , Ting Yan , Ji Zhu

We propose a novel statistical model for sparse networks with overlapping community structure. The model is based on representing the graph as an exchangeable point process, and naturally generalizes existing probabilistic models with…

Methodology · Statistics 2025-02-06 Adrien Todeschini , Xenia Miscouridou , François Caron

Multi-source and multi-modal datasets are increasingly common in scientific research, yet they often exhibit block-wise missingness, where entire modalities are systematically absent in some sources or no single source contains all…

Methodology · Statistics 2026-02-10 Kejian Zhang , Muxuan Liang , Robert Maile , Doudou Zhou

This paper is concerned with the problem of comparing the population means of two groups of independent observations. An approximate randomization test procedure based on the test statistic of Chen and Qin (2010) is proposed. The asymptotic…

Statistics Theory · Mathematics 2022-08-23 Rui Wang , Wangli Xu

A common disadvantage in existing distribution-free two-sample testing approaches is that the computational complexity could be high. Specifically, if the sample size is $N$, the computational complexity of those two-sample tests is at…

Methodology · Statistics 2017-07-18 Cheng Huang , Xiaoming Huo

A number of applications require two-sample testing on ranked preference data. For instance, in crowdsourcing, there is a long-standing question of whether pairwise comparison data provided by people is distributed similar to…

Machine Learning · Statistics 2020-11-20 Charvi Rastogi , Sivaraman Balakrishnan , Nihar B. Shah , Aarti Singh

This paper investigates a statistical procedure for testing the equality of two independent estimated covariance matrices when the number of potentially dependent data vectors is large and proportional to the size of the vectors, that is,…

Statistics Theory · Mathematics 2020-06-01 Rémy Mariétan , Stephan Morgenthaler

Robust classification algorithms have been developed in recent years with great success. We take advantage of this development and recast the classical two-sample test problem in the framework of classification. Based on the estimates of…

Statistics Theory · Mathematics 2019-09-18 Haiyan Cai , Bryan Goggin , Qingtang Jiang

We study the problem of testing for structure in networks using relations between the observed frequencies of small subgraphs. We consider the statistics \begin{align*} T_3 & =(\text{edge frequency})^3 - \text{triangle frequency}\\ T_2 &…

Methodology · Statistics 2017-04-25 Chao Gao , John Lafferty

Networks arise naturally in many scientific fields as a representation of pairwise connections. Statistical network analysis has most often considered a single large network, but it is common in a number of applications to observe multiple…

Methodology · Statistics 2026-03-16 Peter W. MacDonald , Elizaveta Levina , Ji Zhu

We introduce fully nonparametric two-sample tests for testing the null hypothesis that the samples come from the same distribution if the values are only indirectly given via current status censoring. The tests are based on the likelihood…

Statistics Theory · Mathematics 2013-07-12 Piet Groeneboom