Related papers: Selected Methods for non-Gaussian Data Analysis
High dimensional and heterogeneous count data are collected in various applied fields. In this paper, we look closely at high-resolution sequencing data on the microbiome, which have enabled researchers to study the genomes of entire…
The analytic inference, e.g. predictive distribution being in closed form, may be an appealing benefit for machine learning practitioners when they treat wide neural networks as Gaussian process in Bayesian setting. The realistic widths,…
We consider the analysis of continuous repeated measurement outcomes that are collected through time, also known as longitudinal data. A standard framework for analysing data of this kind is a linear Gaussian mixed-effects model within…
The bivariate copulas that describe the dependencies and partial dependencies of lagged variables in strictly stationary, first-order GARCH-type processes are investigated. It is shown that the copulas of symmetric GARCH processes are…
By considering a nonlinear combination of the probabilities of a system, a physical interpretation of Tsallis statistics as representing the nonlinear coupling or decoupling of statistical states is proposed. The escort probability is…
When scholars study joint distributions of multiple variables, copulas are useful. However, if the variables are not linearly correlated with each other yet are still not independent, most of conventional copulas are not up to the task.…
In modern science, computer models are often used to understand complex phenomena, and a thriving statistical community has grown around analyzing them. This review aims to bring a spotlight to the growing prevalence of stochastic computer…
The Gaussian process (GP) is a popular statistical technique for stochastic function approximation and uncertainty quantification from data. GPs have been adopted into the realm of machine learning in the last two decades because of their…
Decentralized learning enables a group of collaborative agents to learn models using a distributed dataset without the need for a central parameter server. Recently, decentralized learning algorithms have demonstrated state-of-the-art…
In some areas of knowledge there are data representing directions restricted to a specific range of values. Consequently, it is useful to have models for describing variables defined in subsets of the k-dimensional unit sphere. This need…
The paradigm of linear structural equation modeling readily allows one to incorporate causal feedback loops in the model specification. These appear as directed cycles in the common graphical representation of the models. However, the…
Modern information processing relies on the axiom that high-dimensional data lie near low-dimensional geometric structures. This paper revisits the problem of data-driven learning of these geometric structures and puts forth two new…
Second-order characteristics including covariance and spectral density functions are fundamentally important for both statistical applications and theoretical analysis in functional time series. In the high-dimensional setting where the…
Quantifying uncertainty in future climate projections is hindered by the prohibitive computational cost of running physical climate models, which severely limits the availability of training data. We propose a data-efficient framework for…
In this paper, we study the identifiability and the estimation of the parameters of a copula-based multivariate model when the margins are unknown and are arbitrary, meaning that they can be continuous, discrete, or mixtures of continuous…
Functional principal component analysis is essential in functional data analysis, but the inferences will become unconvincing when some non-Gaussian characteristics occur, such as heavy tail and skewness. The focus of this paper is to…
Neural networks excel at discovering statistical patterns in high-dimensional data sets. In practice, higher-order cumulants, which quantify the non-Gaussian correlations between three or more variables, are particularly important for the…
We develop a class of nearest-neighbor mixture models that provide direct, computationally efficient, probabilistic modeling for non-Gaussian geospatial data. The class is defined over a directed acyclic graph, which implies conditional…
We generalize the maximum likelihood method to non-Gaussian distribution functions by means of the multivariate Edgeworth expansion. We stress the potential interest of this technique in all those cosmological problems in which the…
Information theoretic measures (entropies, entropy rates, mutual information) are nowadays commonly used in statistical signal processing for real-world data analysis. The present work proposes the use of Auto Mutual Information (Mutual…