English
Related papers

Related papers: A new quantity for statistical analysis: "Scaling …

200 papers

Consider a panel data setting where repeated observations on individuals are available. Often it is reasonable to assume that there exist groups of individuals that share similar effects of observed characteristics, but the grouping is…

Methodology · Statistics 2024-02-09 Lu Yu , Jiaying Gu , Stanislav Volgushev

Approximate Bayesian computation performs approximate inference for models where likelihood computations are expensive or impossible. Instead simulations from the model are performed for various parameter values and accepted if they are…

Computation · Statistics 2015-12-16 Dennis Prangle

In the finite-size scaling analysis of Monte Carlo data, instead of computing the observables at fixed Hamiltonian parameters, one may choose to keep a renormalization-group invariant quantity, also called phenomenological coupling, fixed…

Statistical Mechanics · Physics 2011-08-31 Francesco Parisen Toldin

Statistical modeling is often used to measure the strength of evidence for or against hypotheses on given data. We have previously proposed an information-dynamic framework in support of a properly calibrated measurement scale for…

Statistics Theory · Mathematics 2023-07-19 V. J Vieland , S-C. Seok

From longitudinal biomedical studies to social networks, graphs have emerged as a powerful framework for describing evolving interactions between agents in complex systems. In such studies, after pre-processing, the data can be represented…

Applications · Statistics 2018-03-12 Claire Donnat , Susan Holmes

We study the susceptibility, i.e., the mean size of the component containing a random vertex, in a general model of inhomogeneous random graphs. This is one of the fundamental quantities associated to (percolation) phase transitions; in…

Probability · Mathematics 2012-03-27 Svante Janson , Oliver Riordan

The occurrence of the nonzero leftmost digit, i.e., 1, 2, ..., 9, of numbers from many real world sources is not uniformly distributed as one might naively expect, but instead, the nature favors smaller ones according to a logarithmic…

Data Analysis, Statistics and Probability · Physics 2014-11-21 Lijing Shao , Bo-Qiang Ma

Benford's Law predicts that the first significant digit on the leftmost side of numbers in real-life data is proportioned between all possible 1 to 9 digits approximately as in LOG(1 + 1/digit), so that low digits occur much more frequently…

Physics and Society · Physics 2020-01-22 Alex Ely Kossovsky

Benford's law is the statement that in many real world data sets, the probability of having digit $d$ in base $B$ as the first digit is \log_{B}\!\left(\frac{d+1}{d}\right) for all $1 \leq d \leq B$. We sometimes refer to this as weak…

Probability · Mathematics 2026-03-06 Bruce Fang , Steven J. Miller

Building on recent work in statistical science, the paper presents a theory for modelling natural phenomena that unifies physical and statistical paradigms based on the underlying principle that a model must be nondimensionalizable. After…

Statistics Theory · Mathematics 2021-09-07 Tae Yoon Lee , James V. Zidek , Nancy Heckman

We consider the statistical analysis of data on high-dimensional spheres and shape spaces. The work is of particular relevance to applications where high-dimensional data are available--a commonly encountered situation in many disciplines.…

Statistics Theory · Mathematics 2007-06-13 Ian L. Dryden

(To appear in The American Statistician.) Distance covariance (Sz\'ekely, Rizzo, and Bakirov, 2007) is a fascinating recent notion, which is popular as a test for dependence of any type between random variables $X$ and $Y$. This approach…

Methodology · Statistics 2024-07-08 Jakob Raymaekers , Peter J. Rousseeuw

Diffusion state distance (DSD) is a metric on the vertices of a graph, motivated by bioinformatic modeling. Previous results on the convergence of DSD to a limiting metric relied on the definition being based on symmetric or reversible…

Probability · Mathematics 2015-02-26 Neal Madras

This paper introduces a new framework for recovering causal graphs from observational data, leveraging the observation that the distribution of an effect, conditioned on its causes, remains invariant to changes in the prior distribution of…

Machine Learning · Computer Science 2026-02-04 Nang Hung Nguyen , Phi Le Nguyen , Thao Nguyen Truong , Trong Nghia Hoang , Masashi Sugiyama

Studies on distribution, abundance and diversity of species revealed fascinating universalities in macroecology. Many of these patterns, like the species-area and range-abundance relationship or the year-to-year fluctuations in population…

Populations and Evolution · Quantitative Biology 2007-05-23 M. Ravasz , A. Balog , V. Marko , Z. Neda

Acyclic digraphs arise in many natural and artificial processes. Among the broader set, dynamic citation networks represent a substantively important form of acyclic digraphs. For example, the study of such networks includes the spread of…

Physics and Society · Physics 2011-07-26 Michael J. Bommarito , Daniel Martin Katz , Jon Zelner , James H. Fowler

We introduce the bivariate unit-log-symmetric model based on the bivariate log-symmetric distribution (BLS) defined in [Vila et al., 2022, Bivariate Log-symmetric Models: Theoretical Properties and Parameter Estimation. Avaliable at…

Methodology · Statistics 2023-01-19 Roberto Vila , Narayanaswamy Balakrishnan , Helton Saulo , Peter Zörnig

Distribution shifts, where statistical properties differ between training and test datasets, present a significant challenge in real-world machine learning applications where they directly impact model generalization and robustness. In this…

Machine Learning · Computer Science 2024-05-06 Vegard Flovik

Benford's law is the statement that in many real-world data sets, the probability of having digit \(d\) in base \(B\), where \(1 \leq d \leq B\), as the first digit is \(\log_{B}\left(\tfrac{d+1}{d}\right)\). We sometimes refer to this as…

Probability · Mathematics 2025-08-26 Bruce Fang , Ava Irons , Ella Lippelman , Steven J. Miller

There are many distance-based methods for classification and clustering, and for data with a high number of dimensions and a lower number of observations, processing distances is computationally advantageous compared to the raw data matrix.…

Methodology · Statistics 2020-06-25 Christian Hennig