Related papers: Fitting heavy tailed distributions: the poweRlaw p…
Random forest regression (RF) is an extremely popular tool for the analysis of high-dimensional data. Nonetheless, its benefits may be lessened in sparse settings due to weak predictors, and a pre-estimation dimension reduction (targeting)…
As alternatives to the normal distributions, $t$ distributions are widely applied in robust analysis for data with outliers or heavy tails. The properties of the multivariate $t$ distribution are well documented in Kotz and Nadarajah's…
Logarithmic transformation of the data has been recommended by the literature in the case of highly skewed distributions such as those commonly found in information science. The purpose of the transformation is to make the data conform to…
For a large class of statistical systems a geometric mean value of the observables is constrained. These observables are characterized by a power-law statistical distribution.
Given an arbitrary continuous probability density function, it is introduced a conjugated probability density, which is defined through the Shannon information associated with its cumulative distribution function. These new densities are…
Compute-optimal scaling laws are relatively well studied for NLP and CV, where objectives are typically single-step and targets are comparatively homogeneous. Weather forecasting is harder to characterize in the same framework:…
Modeling distributions of citations to scientific papers is crucial for understanding how science develops. However, there is a considerable empirical controversy on which statistical model fits the citation distributions best. This paper…
R-MAT is a simple, widely used recursive model for generating `complex network' graphs with a power law degree distribution and community structure. We make R-MAT even more useful by reducing the required work per edge from logarithmic to…
Skewness and kurtosis are fundamental statistical moments commonly used to quantify asymmetry and tail behavior in probability distributions. Despite their widespread application in statistical mechanics, condensed matter physics, and…
In risk analysis, a global fit that appropriately captures the body and the tail of the distribution of losses is essential. Modelling the whole range of the losses using a standard distribution is usually very hard and often impossible due…
Heavy-tailed impact distributions, intrinsic uncertainty, and the high costs of proposal-based peer review increasingly challenge research funding decisions. Using large-scale bibliometric data, we show that past scientific performance…
Power-law distributions contain precious information about a large variety of processes in geoscience and elsewhere. Although there are sound theoretical grounds for these distributions, the empirical evidence in favor of power laws has…
It is generally recognized that economical systems, and more in general complex systems, are characterized by power law distributions. Sometime, these distributions show a changing of the slope in the tail so that, more appropriately, they…
The study of heavy-tailed distributions in economic and financial systems has been widely addressed since financial time series has become a research subject.After the eighties, several "highly improbable" market drops were observed (e.g.…
Random forests are an ensemble method relevant for many problems, such as regression or classification. They are popular due to their good predictive performance (compared to, e.g., decision trees) requiring only minimal tuning of…
We develop and implement a version of the popular "policytree" method (Athey and Wager, 2021) using discrete optimisation techniques. We test the performance of our algorithm in finite samples and find an improvement in the runtime of…
When the probability of measuring a particular value of some quantity varies inversely as a power of that value, the quantity is said to follow a power law, also known variously as Zipf's law or the Pareto distribution. Power laws appear…
Many data distributions in the real world are hardly uniform. Instead, skewed and long-tailed distributions of various kinds are commonly observed. This poses an interesting problem for machine learning, where most algorithms assume or work…
Bipartite projections are used in a wide range of network contexts including politics (bill co-sponsorship), genetics (gene co-expression), economics (executive board co-membership), and innovation (patent co-authorship). However, because…
Most random graph models are locally tree-like - do not contain short cycles - rendering them unfit for modeling networks with a community structure. We introduce the hierarchical configuration model (HCM), a generalization of the…