Related papers: A New Approach to Determine the Coefficient of Ske…
This paper presents a new similarity measure to be used for general tasks including supervised learning, which is represented by the K-nearest neighbor classifier (KNN). The proposed similarity measure is invariant to large differences in…
Applications in data science, shape analysis and object classification frequently require comparison of probability distributions defined on different ambient spaces. To accomplish this, one requires a notion of distance on a given class of…
In this paper, the development of a mathematical method is presented to explore spatially non-uniform phases with no long-range order in mathematical models of first order phase transitions. We use essential results regarding the…
In this paper, we develop a local rank correlation measure which quantifies the performance of dimension reduction methods. The local rank correlation is easily interpretable, and robust against the extreme skewness of nearest neighbor…
The asymmetric objective function is proposed as an alternative to Huber objective function to model skewness and obtain robust estimators for the location, scale and skewness parameters. The robustness and asymptotic properties of the…
Kink model is developed to analyze the data where the regression function is twostage linear but intersects at an unknown threshold. In quantile regression with longitudinal data, previous work assumed that the unknown threshold parameters…
Analytical expressions for covariances of weak lensing statistics related to the aperture mass $\Map$ are derived for realistic survey geometries such as SNAP for a range of smoothing angles and redshift bins. We incorporate the…
In this paper, we consider the problem of estimating the distance between any two large data streams in small- space constraint. This problem is of utmost importance in data intensive monitoring applications where input streams are…
Increasingly, critical decisions in public policy, governance, and business strategy rely on a deeper understanding of the needs and opinions of constituent members (e.g. citizens, shareholders). While it has become easier to collect a…
We propose a simple yet powerful test statistic to quantify the discrepancy between two conditional distributions. The new statistic avoids the explicit estimation of the underlying distributions in highdimensional space and it operates on…
The inverse tangent function can be bounded by different inequalities, for example by Shafer's inequality. In this publication, we propose a new sharp double inequality, consisting of a lower and an upper bound, for the inverse tangent…
While recent work has established divergence as a key framework for understanding evenness, there is currently no research exploring how the families of measures within the divergence-based framework relate to each other. This paper uses…
Silhouette coefficient is an established internal clustering evaluation measure that produces a score per data point, assessing the quality of its clustering assignment. To assess the quality of the clustering of the whole dataset, the…
In this paper, we describe a new vector similarity measure associated with a convex cost function. Given two vectors, we determine the surface normals of the convex function at the vectors. The angle between the two surface normals is the…
We propose a fast and scalable algorithm to project a given density on a set of structured measures defined over a compact 2D domain. The measures can be discrete or supported on curves for instance. The proposed principle and algorithm are…
We introduce a quantitative method to compare arbitrary pairs of graph centrality measures, based on the ordering of vertices induced by them. The proposed method is conceptually simple, mathematically elegant, and allows for a quantitative…
On the basis of an analysis of previous research, we present a generalized approach for measuring the difference of plans with an exemplary application to machine scheduling. Our work is motivated by the need for such measures, which are…
Maximal correlation is a measure of correlation for bipartite distributions. This measure has two intriguing features: (1) it is monotone under local stochastic maps; (2) it gives the same number when computed on i.i.d. copies of a pair of…
We study the problem of estimating a manifold from random samples. In particular, we consider piecewise constant and piecewise linear estimators induced by k-means and k-flats, and analyze their performance. We extend previous results for…
This paper investigates the distribution of marks obtained by students across multiple courses to explore whether the data conforms to a skew-normal distribution. Traditional methods for assessing normality, such as the Shapiro Wilk test,…