Related papers: Complexity of hierarchical ensembles
Complexity is a multi-faceted phenomenon, involving a variety of features including disorder, nonlinearity, and self-organisation. We use a recently developed rigorous framework for complexity to understand measures of complexity. We…
A probabilistic database with attribute-level uncertainty consists of relations where cells of some attributes may hold probability distributions rather than deterministic content. Such databases arise, implicitly or explicitly, in the…
G. Edelman, O. Sporns, and G. Tononi have introduced the neural complexity of a family of random variables, defining it as a specific average of mutual information over subfamilies. We show that their choice of weights satisfies two natural…
We investigate the relationship between complexity, information transfer and the emergence of collective behaviors, such as synchronization and nontrivial collective behavior, in a network of globally coupled chaotic maps as a simple model…
Assessing variability according to distinct factors in data is a fundamental technique of statistics. The method commonly regarded to as analysis of variance (ANOVA) is, however, typically confined to the case where all levels of a factor…
Tree ensembles have demonstrated state-of-the-art predictive performance across a wide range of problems involving tabular data. Nevertheless, the black-box nature of tree ensembles is a strong limitation, especially for applications with…
A directly measurable correlation length may be defined for systems having a two-step relaxation, based on the geometric properties of density profile that remains after averaging out the fast motion. We argue that the length diverges if…
Entropy measures have become increasingly popular as an evaluation metric for complexity in the analysis of time series data, especially in physiology and medicine. Entropy measures the rate of information gain, or degree of regularity in a…
We propose a novel "tree-averaging" model that utilizes the ensemble of classification and regression trees (CART). Each constituent tree is estimated with a subset of similar data. We treat this grouping of subsets as Bayesian ensemble…
Complex systems are found in most branches of science. It is still argued how to best quantify their complexity and to what end. One prominent measure of complexity (the statistical complexity) has an operational meaning in terms of the…
The benefits of diversifying risks are difficult to estimate quantitatively because of the uncertainties in the dependence structure between the risks. Also, the modelling of multidimensional dependencies is a non-trivial task. This paper…
Graphs are used to represent and analyze data in domains as diverse as physics, biology, chemistry, planetary science, and the social sciences. Across domains, random graph models relate generative processes to expected graph properties,…
An axiomatic foundation regarding the entropy for complex systems is established. Missing from decades of research was the requirement that entropy must measure the uncertainty at the informational scale of the maximizing distribution,…
The ETH ansatz for matrix elements of a given operator in the energy eigenstate basis results in a notion of thermalization for a chaotic system. In this context for a certain quantity - to be found for a given model - one may impose a…
Probability estimation is one of the fundamental tasks in statistics and machine learning. However, standard methods for probability estimation on discrete objects do not handle object structure in a satisfactory manner. In this paper, we…
A non-ergodic quantum state of a many body system is in general random as well as multi-parametric, former due to a lack of exact information due to complexity and latter reflecting its varied behavior in different parts of the Hilbert…
In ensemble methods, the outputs of a collection of diverse classifiers are combined in the expectation that the global prediction be more accurate than the individual ones. Heterogeneous ensembles consist of predictors of different types,…
We consider a Gaussian statistical model whose parameter space is given by the variances of random variables. Underlying this model we identify networks by interpreting random variables as sitting on vertices and their correlations as…
Tree ensembles (TEs) find a multitude of practical applications. They represent one of the most general and accurate classes of machine learning methods. While they are typically quite concise in representation, their operation remains…
Information theory on a time-discrete setting in the framework of time series analysis is generalized to the time-continuous case. Considerations of the Roessler and Lorenz dynamics as well as the Ornstein-Uhlenbeck process yield for…