English
Related papers

Related papers: Adjusting for Chance Clustering Comparison Measure…

200 papers

Rare events attract more attention and interests in many scenarios of big data such as anomaly detection and security systems. To characterize the rare events importance from probabilistic perspective, the message importance measure (MIM)…

Information Theory · Computer Science 2019-01-07 Rui She , Shanyun Liu , Pingyi Fan

In this paper, we provide an approach to clustering relational matrices whose entries correspond to either similarities or dissimilarities between objects. Our approach is based on the value of information, a parameterized,…

Artificial Intelligence · Computer Science 2017-10-31 Isaac J. Sledge , Jose C. Principe

We propose a penalized likelihood method to jointly estimate multiple precision matrices for use in quadratic discriminant analysis and model based clustering. A ridge penalty and a ridge fusion penalty are used to introduce shrinkage and…

Machine Learning · Statistics 2014-05-06 Bradley S. Price , Charles J. Geyer , Adam J. Rothman

For the use case of comparing the performance of clustering algorithms whose output is a contingency table, a single performance metric for contingency tables is needed. Such a metric is vital for comparative performance analysis of…

Machine Learning · Computer Science 2026-05-01 Naomi E. Zirkind , William J. Diehl

Determining the number of common factors is an important and practical topic in high dimensional factor models. The existing literatures are mainly based on the eigenvalues of the covariance matrix. Due to the incomparability of the…

Methodology · Statistics 2019-09-25 Jianqing Fan , Jianhua Guo , Shurong Zheng

Information-theoretic quantities like entropy and mutual information have found numerous uses in machine learning. It is well known that there is a strong connection between these entropic quantities and submodularity since entropy over a…

Machine Learning · Computer Science 2021-03-04 Rishabh Iyer , Ninad Khargonkar , Jeff Bilmes , Himanshu Asnani

Subtraction of aligned images is a means to assess changes in a wide variety of clinical applications. In this paper we explore the information theoretical origin of Mutual Information (MI), which is based on Shannon's entropy.However, the…

Computer Vision and Pattern Recognition · Computer Science 2007-06-19 W. Jacquet , P. de Groen

More than twenty years after its introduction, Annealed Importance Sampling (AIS) remains one of the most effective methods for marginal likelihood estimation. It relies on a sequence of distributions interpolating between a tractable…

Machine Learning · Statistics 2022-10-25 Arnaud Doucet , Will Grathwohl , Alexander G. D. G. Matthews , Heiko Strathmann

The Adaptive Multiple Importance Sampling (AMIS) algorithm is aimed at an optimal recycling of past simulations in an iterated importance sampling scheme. The difference with earlier adaptive importance sampling implementations like…

Computation · Statistics 2011-10-04 Jean-Marie Cornuet , Jean-Michel Marin , Antonietta Mira , Christian P. Robert

Multivariate time-series anomaly detection (MTSAD) aims to identify deviations from normality in multivariate time-series and is critical in real-world applications. However, in real-world deployments, distribution shifts are ubiquitous and…

Machine Learning · Computer Science 2026-04-03 HyunGi Kim , Jisoo Mok , Hyungyu Lee , Juhyeon Shin , Sungroh Yoon

Covariate-adaptive randomization (CAR) procedures are frequently used in comparative studies to increase the covariate balance across treatment groups. However, because randomization inevitably uses the covariate information when forming…

Statistics Theory · Mathematics 2022-07-08 Wei Ma , Yichen Qin , Yang Li , Feifang Hu

For better clustering performance, appropriate representations are critical. Although many neural network-based metric learning methods have been proposed, they do not directly train neural networks to improve clustering performance. We…

Machine Learning · Statistics 2021-03-02 Tomoharu Iwata

Generalized mutual information (GMI) is used to compute achievable rates for fading channels with various types of channel state information at the transmitter (CSIT) and receiver (CSIR). The GMI is based on variations of auxiliary channel…

Information Theory · Computer Science 2023-05-23 Gerhard Kramer

Entropy and relative or cross entropy measures are two very fundamental concepts in information theory and are also widely used for statistical inference across disciplines. The related optimization problems, in particular the maximization…

Statistics Theory · Mathematics 2021-06-18 Abhik Ghosh , Ayanendranath Basu

Tsallis and R\'{e}nyi entropy measures are two possible different generalizations of the Boltzmann-Gibbs entropy (or Shannon's information) but are not generalizations of each others. It is however the Sharma-Mittal measure, which was…

Statistical Mechanics · Physics 2014-10-13 Marco Masi

This paper proposes a new method for estimating the joint probability mass function of a pair of discrete random variables. This estimator is used to construct joint Shannon R\'enyi-Tsallis entropies, and the mutual information estimates of…

Methodology · Statistics 2020-01-14 Amadou Diadie Ba , Gane Samb Lo , Cheikh Tidiane Seck

Performance of clustering algorithms is evaluated with the help of accuracy metrics. There is a great diversity of clustering algorithms, which are key components of many data analysis and exploration systems. However, there exist only few…

Data Structures and Algorithms · Computer Science 2019-02-18 Artem Lutov , Mourad Khayati , Philippe Cudré-Mauroux

In this article, we discuss the problem of establishing relations between information measures assessed for network structures. Two types of entropy based measures namely, the Shannon entropy and its generalization, the R\'{e}nyi entropy…

Information Theory · Computer Science 2013-01-24 Lavanya Sivakumar , Matthias Dehmer

Prediction-Powered Inference (PPI) is a powerful framework for enhancing statistical estimates by combining limited gold-standard data with machine learning (ML) predictions. While prior work has demonstrated PPI's benefits for individual…

Machine Learning · Statistics 2025-11-10 Sida Li , Nikolaos Ignatiadis

In electronic health records (EHR) analysis, clustering patients according to patterns in their data is crucial for uncovering new subtypes of diseases. Existing medical literature often relies on classical hypothesis testing methods to…

Methodology · Statistics 2024-05-07 Zihan Zhu , Xin Gai , Anru R. Zhang