Related papers: Estimation of Goodness-of-Fit in Multidimensional …
We study the Bahadur efficiency of several weighted L2--type goodness--of--fit tests based on the empirical characteristic function. The methods considered are for normality and exponentiality testing, and for testing goodness--of--fit to…
We give a general unified method that can be used for $L_1$ {\em closeness testing} of a wide range of univariate structured distribution families. More specifically, we design a sample optimal and computationally efficient algorithm for…
Single-cell omics enable the profiles of cells, which contain large numbers of biological features, to be quantified. Cluster analysis, a dimensionality reduction process, is used to reduce the dimensions of the data to make it…
In this paper a relative number density parameter, called the neighborhood function, is introduced so that the crowded nature of the neighborhood of individual sources can be described. With this parameter one can determine the probability…
Most of the existing methods for estimating the local intrinsic dimension of a data distribution do not scale well to high-dimensional data. Many of them rely on a non-parametric nearest neighbors approach which suffers from the curse of…
Using the fact that some depth functions characterize certain family of distribution functions, and under some mild conditions, distribution of the depth is continuous, we have constructed several new multivariate goodness of fit tests…
The chi square goodness-of-fit test is among the oldest known statistical tests, first proposed by Pearson in 1900 for the multinomial distribution. It has been in use in many fields ever since. However, various studies have shown that when…
This article deals with goodness-of-fit test for the Cauchy distribution. Some tests based on Kullback-Leibler information are proposed, and shown to be consistent. Monte Carlo evidence indicates that the tests have satisfactory…
Record is used to reduce the time and cost of running experiments (Doostparast and Balakrishnan, 2010). It is important to check the adequacy of models upon which inferences or actions are based (Lawless, 2003, Chapter 10, p. 465). In the…
Many data mining and statistical machine learning algorithms have been developed to select a subset of covariates to associate with a response variable. Spurious discoveries can easily arise in high-dimensional data analysis due to enormous…
We present a review of several results concerning the construction of the Cramer-von Mises and Kolmogorov-Smirnov type goodness-of-fit tests for continuous time processes. As the models we take a stochastic differential equation with small…
In the context of cluster analysis and graph partitioning, many external evaluation measures have been proposed in the literature to compare two partitions of the same set. This makes the task of selecting the most appropriate measure for a…
Using a sample of $>200$ clusters, each with typically $100-200$ spectroscopically confirmed cluster members, we search for a signal of alignment between the Position Angle (PA) of the Brightest Cluster Galaxy (BCG) and the distribution of…
We propose an empirical likelihood test that is able to test the goodness of fit of a class of parametric and semi-parametric multiresponse regression models. The class includes as special cases fully parametric models; semi-parametric…
We introduce the QuadratiK package that incorporates innovative data analysis methodologies. The presented software, implemented in both R and Python, offers a comprehensive set of goodness-of-fit tests and clustering techniques using…
Identifying the number $K$ of clusters in a dataset is one of the most difficult problems in clustering analysis. A choice of $K$ that correctly characterizes the features of the data is essential for building meaningful clusters. In this…
The new particle accelerators and its experiments create a challenging data processing environment, characterized by large amount of data where only small portion of it carry the expected new scientific information. Modern detectors, such…
Many objects studied in astronomy follow a power law distribution function, for example the masses of stars or star clusters. A still used method by which such data is analysed is to generate a histogram and fit a straight line to it. The…
Tests of goodness of fit are used in nearly every domain where statistics is applied. One powerful and flexible approach is to sample artificial data sets that are exchangeable with the real data under the null hypothesis (but not under the…
Generalized linear models (GLMs) are used within a vast number of application domains. However, formal goodness of fit (GOF) tests for the overall fit of the model$-$so-called "global" tests$-$seem to be in wide use only for certain classes…