English
Related papers

Related papers: Segregation Indices for Disease Clustering

200 papers

We derive the finite size dependence of the clustering coefficient of scale-free random graphs generated by the configuration model with degree distribution exponent $2<\gamma<3$. Degree heterogeneity increases the presence of triangles in…

Disordered Systems and Neural Networks · Physics 2015-06-05 Pol Colomer-de-Simon , Marian Boguna

A recent proposal of data dependent similarity called Isolation Kernel/Similarity has enabled SVM to produce better classification accuracy. We identify shortcomings of using a tree method to implement Isolation Similarity; and propose a…

Machine Learning · Computer Science 2024-01-30 Xiaoyu Qin , Kai Ming Ting , Ye Zhu , Vincent CS Lee

Independence analysis is an indispensable step before regression analysis to find out essential factors that influence the objects. With many applications in machine Learning, medical Learning and a variety of disciplines, statistical…

Methodology · Statistics 2022-07-08 Wenliang Pan , Yujue Li , Jianwu Liu , Pei Dang , Weixiong Mai

Mapping of spatial hotspots, i.e., regions with significantly higher rates of generating cases of certain events (e.g., disease or crime cases), is an important task in diverse societal domains, including public health, public safety,…

Machine Learning · Statistics 2021-10-12 Yiqun Xie , Shashi Shekhar , Yan Li

In this paper, we consider feature screening for ultrahigh dimensional clustering analyses. Based on the observation that the marginal distribution of any given feature is a mixture of its conditional distributions in different clusters, we…

Methodology · Statistics 2024-02-05 Changhu Wang , Zihao Chen , Ruibin Xi

Whenever possible, the efficacy of a new treatment, such as a drug or behavioral intervention, is investigated by randomly assigning some individuals to a treatment condition and others to a control condition, and comparing the outcomes…

Methodology · Statistics 2015-05-04 Patrick C. Staples , Elizabeth L. Ogburn , Jukka-Pekka Onnela

Suppose that we are interested in the comparison of two independent categorical variables. Suppose also that the population is divided into subpopulations or groups. Notice that the distribution of the target variable may vary across…

Methodology · Statistics 2024-05-08 M. V. Alba-Fernández , M. D. Jiménez--Gamero , F. J. Ariza-López

Clustering methods are a valuable tool for the identification of patterns in high dimensional data with applications in many scientific problems. However, quantifying uncertainty in clustering is a challenging problem, particularly when…

Methodology · Statistics 2018-06-01 Marcio Valk , Gabriela Bettella Cybis

Traditionally, the Dirichlet-multinomial distribution has been recognized as a key model for contingency tables generated by cluster sampling schemes. There are, however, other possible distributions appropriate for these contingency…

Methodology · Statistics 2016-09-26 Juana M. Alonso-Revenga , Nirian Martin , Leandro Pardo

In clinical trials studying paired parts of a subject with binary outcomes, it is expected to collect measurements bilaterally. However, there are cases where subjects contribute measurements for only one part. By utilizing combined data,…

Applications · Statistics 2024-03-06 Shuyi Liang , Kai-Tai Fang , Xin-Wei Huang , Yijing Xin , Chang-Xing Ma

A compact metric space $(X, \rho)$ is given. Let $\mu$ be a Borel measure on $X$. By $r$-cluster we mean a measurable subset of $X$ with diameter at most $r$. A family of $k$ $2r$-clusters is called a $r$-cluster structure of order $k$ if…

Discrete Mathematics · Computer Science 2017-09-26 Alexey Pushnyakov

In this paper, we consider the problem of partitioning a small data sample of size $n$ drawn from a mixture of $2$ sub-gaussian distributions. Our work is motivated by the application of clustering individuals according to their population…

Statistics Theory · Mathematics 2023-01-05 Shuheng Zhou

A cluster tree provides a highly-interpretable summary of a density function by representing the hierarchy of its high-density clusters. It is estimated using the empirical tree, which is the cluster tree constructed from a density…

Statistics Theory · Mathematics 2017-02-14 Jisu Kim , Yen-Chi Chen , Sivaraman Balakrishnan , Alessandro Rinaldo , Larry Wasserman

Generalized likelihood ratio (GLR) test statistics are often used in the detection of spatial clustering in case-control and case-population datasets to check for a significantly large proportion of cases within some scanning window. The…

Statistics Theory · Mathematics 2009-11-20 Hock Peng Chan

Urban areas with larger and more connected populations offer an auspicious environment for contagion processes such as the spread of pathogens. Empirical evidence reveals a systematic increase in the rates of certain sexually transmitted…

Physics and Society · Physics 2018-09-18 Oscar Patterson-Lomba , Andres Gomez-Lievano

We investigate the effects of heterogeneous and clustered contact patterns on the timescale and final size of infectious disease epidemics. The abundance of transitive relationships (the number of 3 cliques) in a network and the variance of…

Quantitative Methods · Quantitative Biology 2010-06-07 Erik M Volz

For testing the statistical significance of a treatment effect, we usually compare between two parts of a population, one is exposed to the treatment, and the other is not exposed to it. Standard parametric and nonparametric two-sample…

Computation · Statistics 2012-11-02 Bikram Karmakar , Kumaresh Dhara , Kushal Kumar Dey , Analabha Basu , Anil Ghosh

Cross-level interactions among fixed effects in linear mixed models (also known as multilevel models) are often complicated by the variances stemming from random effects and residuals. When these variances change across clusters, tests of…

Methodology · Statistics 2022-03-18 Ting Wang , Edgar C. Merkle , Joaquin A. Anguera , Brandon M. Turner

Urban living in modern large cities has significant adverse effects on health, increasing the risk of several chronic diseases. We focus on the two leading clusters of chronic disease, heart disease and diabetes, and develop data-driven…

Machine Learning · Computer Science 2018-01-08 Theodora S. Brisimi , Tingting Xu , Taiyao Wang , Wuyang Dai , William G. Adams , Ioannis Ch. Paschalidis

The aim of this Thesis is to present five new tests for random numbers, which are widely used {\em e.g.} in computer simulations in physics applications. The first two tests, the cluster test and the autocorrelation test, are based on…

Condensed Matter · Physics 2008-02-03 I. Vattulainen