English
Related papers

Related papers: Exploring and measuring non-linear correlations: C…

200 papers

In recent biomedical scientific problems, it is a fundamental issue to integratively cluster a set of objects from multiple sources of datasets. Such problems are mostly encountered in genomics, where data is collected from various sources,…

We consider the problem of grouping items into clusters based on few random pairwise comparisons between the items. We introduce three closely related algorithms for this task: a belief propagation algorithm approximating the Bayes optimal…

Social and Information Networks · Computer Science 2016-08-26 Alaa Saade , Marc Lelarge , Florent Krzakala , Lenka Zdeborová

We present a general approach for studying autoregressive categorical time series models with dependence of infinite order and defined conditional on an exogenous covariate process. To this end, we adapt a coupling approach, developed in…

Statistics Theory · Mathematics 2019-08-01 Lionel Truquet

Most data is multi-dimensional. Discovering whether any subset of dimensions, or subspaces, of such data is significantly correlated is a core task in data mining. To do so, we require a measure that quantifies how correlated a subspace is.…

Machine Learning · Statistics 2015-11-12 Hoang-Vu Nguyen , Jilles Vreeken

We propose a novel probabilistic approach to multilevel clustering problems based on composite transportation distance, which is a variant of transportation distance where the underlying metric is Kullback-Leibler divergence. Our method…

Machine Learning · Computer Science 2018-10-30 Nhat Ho , Viet Huynh , Dinh Phung , Michael I. Jordan

In multiple correspondence analysis, both individuals (observations) and categories can be represented in a biplot that jointly depicts the relationships across categories or individuals, as well as the associations between them. Additional…

Methodology · Statistics 2019-01-10 Mariko Takagishi , Michel van de Velden

Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. For most applications, applying clustering is only appropriate when cluster structure is present. As such, the study of clusterability,…

Machine Learning · Statistics 2018-10-30 A. Adolfsson , M. Ackerman , N. C. Brownstein

We propose a clustering procedure to group K populations into subgroups with the same dependence structure. The method is adapted to paired population and can be used with panel data. It relies on the differences between orthogonal…

Methodology · Statistics 2022-11-14 Yves Ismaël Ngounou Bakam , Denys Pommeret

Copulas provide an attractive approach for constructing multivariate distributions with flexible marginal distributions and different forms of dependences. Of particular importance in many areas is the possibility of explicitly forecasting…

Methodology · Statistics 2018-05-22 Feng Li , Yanfei Kang

This paper introduces vector copulas associated with multivariate distributions with given multivariate marginals, based on the theory of measure transportation, and establishes a vector version of Sklar's theorem. The latter provides a…

Econometrics · Economics 2021-04-14 Yanqin Fan , Marc Henry

This paper introduces a nonparametric copula-based index for detecting the strength and monotonicity structure of linear and nonlinear statistical dependence between pairs of random variables or stochastic signals. Our index, termed Copula…

Machine Learning · Statistics 2020-02-25 Kiran Karra , Lamine Mili

An extension of the latent class model is presented for clustering categorical data by relaxing the classical "class conditional independence assumption" of variables. This model consists in grouping the variables into inter-independent and…

Computation · Statistics 2015-10-01 Matthieu Marbac , Christophe Biernacki , Vincent Vandewalle

This paper aims at comparing two coupling approaches as basic layers for building clustering criteria, suited for modularizing and clustering very large networks. We briefly use "optimal transport theory" as a starting point, and a way as…

Discrete Mathematics · Computer Science 2021-03-19 Pierre Bertrand , Michel Broniatowski , Jean-François Marcotorchino

Change point analysis has applications in a wide variety of fields. The general problem concerns the inference of a change in distribution for a set of time-ordered observations. Sequential detection is an online version in which new data…

Methodology · Statistics 2013-10-16 David S. Matteson , Nicholas A. James

In situations where both extreme and non-extreme data are of interest, modelling the whole data set accurately is important. In a univariate framework, modelling the bulk and tail of a distribution has been extensively studied before.…

Methodology · Statistics 2023-10-11 Lídia M. André , Jennifer L. Wadsworth , Adrian O'Hagan

Describing the complex dependence structure of extreme phenomena is particularly challenging. To tackle this issue we develop a novel statistical algorithm that describes extremal dependence taking advantage of the inherent hierarchical…

Methodology · Statistics 2018-07-24 Sabrina Vettori , Raphaël Huser , Johan Segers , Marc G. Genton

Recently, clustering moving object trajectories kept gaining interest from both the data mining and machine learning communities. This problem, however, was studied mainly and extensively in the setting where moving objects can move freely…

Machine Learning · Statistics 2015-11-05 Mohamed Khalil El Mahrsi , Romain Guigourès , Fabrice Rossi , Marc Boullé

In this paper, we introduce Adaptive Cluster Lasso(ACL) method for variable selection in high dimensional sparse regression models with strongly correlated variables. To handle correlated variables, the concept of clustering or grouping…

Machine Learning · Statistics 2016-03-14 Niharika Gauraha , Swapan K. Parui

The R package pdfCluster performs cluster analysis based on a nonparametric estimate of the density of the observed variables. After summarizing the main aspects of the methodology, we describe the features and the usage of the package, and…

Computation · Statistics 2013-01-29 Adelchi Azzalini , Giovanna Menardi

Conditional copulas are flexible statistical tools that couple joint conditional and marginal conditional distributions. In a linear regression setting with more than one covariate and two dependent outcomes, we propose the use of additive…

Methodology · Statistics 2014-07-31 Avideh Sabeti , Mian Wei , Radu V. Craiu