English
Related papers

Related papers: A new quantity for statistical analysis: "Scaling …

200 papers

A measure of primal importance for capturing the serial dependence of a stationary time series at extreme levels is provided by the limiting cluster size distribution. New estimators based on a blocks declustering scheme are proposed and…

Statistics Theory · Mathematics 2020-11-11 Axel Bücher , Tobias Jennessen

Symbolic data analysis (SDA) is an emerging area of statistics concerned with understanding and modelling data that takes distributional form (i.e. symbols), such as random lists, intervals and histograms. It was developed under the premise…

Computation · Statistics 2020-04-09 Boris Beranger , Huan Lin , Scott A. Sisson

We propose a new class of metrics on sets, vectors, and functions that can be used in various stages of data mining, including exploratory data analysis, learning, and result interpretation. These new distance functions unify and generalize…

The Jaccard similarity index has often been employed in science and technology as a means to quantify the similarity between two sets. When modified to operate on real-valued values, the Jaccard similarity index can be applied to compare…

Data Analysis, Statistics and Probability · Physics 2024-10-23 Gonzalo Travieso , Alexandre Benatti , Luciano da F. Costa

Consider a finite directed graph without cycles in which the arrows are weighted. We present an algorithm for the computation of a new distance, called path-length-weighted distance, which has proven useful for graph analysis in the context…

Data Structures and Algorithms · Computer Science 2024-11-05 R. Arnau , J. M. Calabuig , L. M. García Raffi , E. A. Sánchez Pérez , S. Sanjuan

The value of a social network is generally determined by its size and the connectivity of its nodes. But since some of the nodes may be fake ones and others that are dormant, the question of validating the node counts by statistical tests…

Social and Information Networks · Computer Science 2015-06-11 Sieteng Soh , Gongqi Lin , Subhash Kak

Thousands of experiments are analyzed and papers are published each year involving the statistical analysis of grouped data. While this area of statistics is often perceived -- somewhat naively -- as saturated, several misconceptions still…

Methodology · Statistics 2026-02-02 Sara Algeri , Estate V. Khmaladze

Nature and our world have a bias! Roughly $30\%$ of the time the number $1$ occurs as the leading digit in many datasets base $10$. This phenomenon is known as Benford's law and it arrises in diverse fields such as the stock market,…

Probability · Mathematics 2023-08-16 Irfan Durmić , Steven J. Miller

In the framework of Symbolic Data Analysis (SDA), distribution-variables are a particular case of multi-valued variables: each unit is represented by a set of distributions (e.g. histograms, density functions or quantile functions), one for…

Methodology · Statistics 2018-04-20 Rosanna Verde , Antonio Irpino

Linear rate equations are used to describe the cascading decay of an initial heavy cluster into fragments. We consider moments of arbitrary orders of the mass multiplicity spectrum and derive scaling properties pertaining to their time…

Nuclear Theory · Physics 2008-11-26 B. G. Giraud , R. Peschanski

Multivariate time series are ubiquitous objects in signal processing. Measuring a distance or similarity between two such objects is of prime interest in a variety of applications, including machine learning, but can be very difficult as…

Machine Learning · Statistics 2022-11-02 Titouan Vayer , Romain Tavenard , Laetitia Chapel , Nicolas Courty , Rémi Flamary , Yann Soullard

In this paper, a new measurement to compare two large-scale graphs based on the theory of quantum probability is proposed. An explicit form for the spectral distribution of the corresponding adjacency matrix of a graph is established. Our…

Discrete Mathematics · Computer Science 2018-07-03 Hayoung Choi , Hosoo Lee , Yifei Shen , Yuanming Shi

Big Data involves both a large number of events but also many variables. This paper will concentrate on the challenge presented by the large number of variables in a Big Dataset. It will start with a brief review of exploratory data…

Applications · Statistics 2019-07-24 S. J. Watts , L. Crow

Structural properties of evolving random graphs are investigated. Treating linking as a dynamic aggregation process, rate equations for the distribution of node to node distances (paths) and of cycles are formulated and solved analytically.…

Statistical Mechanics · Physics 2007-05-23 E. Ben-Naim , P. L. Krapivsky

A new approach is presented to describe the change in the statistics of the log return distribution of financial data as a function of the timescale. To this purpose a measure is introduced, which quantifies the distance of a considered…

Data Analysis, Statistics and Probability · Physics 2009-11-11 Andreas P. Nawroth , Joachim Peinke

Accurate estimation for extent of cross{sectional dependence in large panel data analysis is paramount to further statistical analysis on the data under study. Grouping more data with weak relations (cross{sectional dependence) together…

Econometrics · Economics 2019-04-16 Jiti Gao , Guangming Pan , Yanrong Yang , Bo Zhang

Let $S$ be a finite set, and $X_1,\ldots,X_n$ an i.i.d. uniform sample from $S$. To estimate the size $|S|$, without further structure, one can wait for repeats and use the birthday problem. This requires a sample size of the order…

Statistics Theory · Mathematics 2026-04-28 Sourav Chatterjee , Persi Diaconis , Susan Holmes

Symbolic Data Analysis (SDA) is a relatively new field of statistics that extends conventional data analysis by taking into account intrinsic data variability and structure. Unlike conventional data analysis, in SDA the features…

Statistics Theory · Mathematics 2021-01-27 M. Rosário Oliveira , Margarida Azeitona , António Pacheco , Rui Valadas

Conventional statistics begins with a model, and assigns a likelihood of obtaining any particular set of data. The opposite approach, beginning with the data and assigning a likelihood to any particular model, is explored here for the case…

Data Analysis, Statistics and Probability · Physics 2009-10-30 Timothy E. Holy

We introduce the notion of symmetric covariation, which is a new measure of dependence between two components of a symmetric $\alpha$-stable random vector, where the stability parameter $\alpha$ measures the heavy-tailedness of its…

Statistics Theory · Mathematics 2021-05-20 Yujia Ding , Qidi Peng