English
Related papers

Related papers: Adjusting for Chance Clustering Comparison Measure…

200 papers

The data for many classification problems, such as pattern and speech recognition, follow mixture distributions. To quantify the optimum performance for classification tasks, the Shannon mutual information is a natural information-theoretic…

Signal Processing · Electrical Eng. & Systems 2022-06-22 Yijun Ding , Amit Ashok

Mutual information (MI) is a fundamental quantity in information theory and machine learning. However, direct estimation of MI is intractable, even if the true joint probability density for the variables of interest is known, as it involves…

Machine Learning · Computer Science 2024-04-29 Rob Brekelmans , Sicong Huang , Marzyeh Ghassemi , Greg Ver Steeg , Roger Grosse , Alireza Makhzani

The use of mutual information as a similarity measure in agglomerative hierarchical clustering (AHC) raises an important issue: some correction needs to be applied for the dimensionality of variables. In this work, we formulate the decision…

Machine Learning · Statistics 2016-08-07 Guillaume Marrelec , Arnaud Messé , Pierre Bellec

This paper studies inference in cluster randomized trials where treatment status is determined according to a "matched pairs" design. Here, by a cluster randomized experiment, we mean one in which treatment is assigned at the level of the…

Econometrics · Economics 2025-08-14 Yuehao Bai , Jizhou Liu , Azeem M. Shaikh , Max Tabord-Meehan

There has been a wide interest to extend univariate and multivariate nonparametric procedures to clustered and hierarchical data. Traditionally, parametric mixed models have been used to account for the correlation structures among the…

Statistics Theory · Mathematics 2018-03-02 Jaakko Nevalainen , Denis Larocque , Hannu Oja , Ilkka Pörsti

Composite likelihood inference has gained much popularity thanks to its computational manageability and its theoretical properties. Unfortunately, performing composite likelihood ratio tests is inconvenient because of their awkward…

Computation · Statistics 2014-08-01 Manuela Cattelan , Nicola Sartori

The analysis of execution paths (also known as software traces) collected from a given software product can help in a number of areas including software testing, software maintenance and program comprehension. The lack of a scalable…

Software Engineering · Computer Science 2012-05-14 A. V. Miranskyy , M. Davison , M. Reesor , S. S. Murtaza

Clustering is considered a non-supervised learning setting, in which the goal is to partition a collection of data points into disjoint clusters. Often a bound $k$ on the number of clusters is given or assumed by the practitioner. Many…

Machine Learning · Computer Science 2012-02-01 Nir Ailon , Ron Begleiter

Several clustering methods (e.g., Normalized Cut and Ratio Cut) divide the Min Cut cost function by a cluster dependent factor (e.g., the size or the degree of the clusters), in order to yield a more balanced partitioning. We, instead,…

Machine Learning · Computer Science 2025-02-06 Morteza Haghir Chehreghani

The abundance of training data is not guaranteed in various supervised learning applications. One of these situations is the post-earthquake regional damage assessment of buildings. Querying the damage label of each building requires a…

Machine Learning · Computer Science 2021-08-17 Mohamadreza Sheibani , Ge Ou

We describe how to calculate standard errors for A/B tests that include clustered data, ratio metrics, and/or covariate adjustment. We may do this for power analysis/sample size calculations prior to running an experiment using historical…

Methodology · Statistics 2024-06-12 Tim Hesterberg , Ben Knight

We study constrained clustering, where constraints guide the clustering process. In existing works, two categories of constraints have been widely explored, namely pairwise and cardinality constraints. Pairwise constraints enforce the…

Machine Learning · Computer Science 2023-01-30 Adel Bibi , Ali Alqahtani , Bernard Ghanem

We introduce the Integrated Tsallis Combination (ITC), a hybrid impurity measure for decision tree learning that combines normalized Tsallis entropy with an exponential polarization component. While many existing measures sacrifice…

Machine Learning · Statistics 2026-03-17 Edouard Lansiaux , Idriss Jairi , Hayfa Zgaya-Biau

There is a growing need for pluralistic alignment methods that can steer language models towards individual attributes and preferences. One such method, Self-Supervised Alignment with Mutual Information (SAMI), uses conditional mutual…

Computation and Language · Computer Science 2025-06-06 Soham V. Govande

In an earlier study, we showed that Tsallis relative entropy (TRE), which is the generalization of Kullback-Leibler relative entropy (KLRE) to non-extensive systems, can be used as a possible risk measure in constructing risk optimal…

Statistical Finance · Quantitative Finance 2022-05-30 Sandhya Devi , Sherman Page

The information theoretic quantity known as mutual information finds wide use in classification and community detection analyses to compare two classifications of the same set of objects into groups. In the context of classification…

Social and Information Networks · Computer Science 2020-04-29 M. E. J. Newman , George T. Cantwell , Jean-Gabriel Young

This paper provides a user's guide to the general theory of approximate randomization tests developed in Canay, Romano, and Shaikh (2017) when specialized to linear regressions with clustered data. An important feature of the methodology is…

Econometrics · Economics 2022-03-16 Yong Cai , Ivan A. Canay , Deborah Kim , Azeem M. Shaikh

Clustering algorithms are used extensively in data analysis for data exploration and discovery. Technological advancements lead to continually growth of data in terms of volume, dimensionality and complexity. This provides great…

Machine Learning · Computer Science 2024-02-20 Miles McCrory , Spencer A. Thomas

We study methods for aggregating pairwise comparison data in order to estimate outcome probabilities for future comparisons among a collection of n items. Working within a flexible framework that imposes only a form of strong stochastic…

Machine Learning · Computer Science 2016-03-23 Nihar B. Shah , Sivaraman Balakrishnan , Martin J. Wainwright

In many health policy settings, adaptive interventions target a population of clusters (e.g., schools), with the ultimate intent of impacting outcomes at the level of individuals within the clusters. Health policy researchers can use…