中文
相关论文

相关论文: Adjusting for Chance Clustering Comparison Measure…

200 篇论文

The data for many classification problems, such as pattern and speech recognition, follow mixture distributions. To quantify the optimum performance for classification tasks, the Shannon mutual information is a natural information-theoretic…

信号处理 · 电气工程与系统科学 2022-06-22 Yijun Ding , Amit Ashok

Mutual information (MI) is a fundamental quantity in information theory and machine learning. However, direct estimation of MI is intractable, even if the true joint probability density for the variables of interest is known, as it involves…

机器学习 · 计算机科学 2024-04-29 Rob Brekelmans , Sicong Huang , Marzyeh Ghassemi , Greg Ver Steeg , Roger Grosse , Alireza Makhzani

The use of mutual information as a similarity measure in agglomerative hierarchical clustering (AHC) raises an important issue: some correction needs to be applied for the dimensionality of variables. In this work, we formulate the decision…

机器学习 · 统计学 2016-08-07 Guillaume Marrelec , Arnaud Messé , Pierre Bellec

This paper studies inference in cluster randomized trials where treatment status is determined according to a "matched pairs" design. Here, by a cluster randomized experiment, we mean one in which treatment is assigned at the level of the…

计量经济学 · 经济学 2025-08-14 Yuehao Bai , Jizhou Liu , Azeem M. Shaikh , Max Tabord-Meehan

There has been a wide interest to extend univariate and multivariate nonparametric procedures to clustered and hierarchical data. Traditionally, parametric mixed models have been used to account for the correlation structures among the…

统计理论 · 数学 2018-03-02 Jaakko Nevalainen , Denis Larocque , Hannu Oja , Ilkka Pörsti

Composite likelihood inference has gained much popularity thanks to its computational manageability and its theoretical properties. Unfortunately, performing composite likelihood ratio tests is inconvenient because of their awkward…

统计计算 · 统计学 2014-08-01 Manuela Cattelan , Nicola Sartori

The analysis of execution paths (also known as software traces) collected from a given software product can help in a number of areas including software testing, software maintenance and program comprehension. The lack of a scalable…

软件工程 · 计算机科学 2012-05-14 A. V. Miranskyy , M. Davison , M. Reesor , S. S. Murtaza

Clustering is considered a non-supervised learning setting, in which the goal is to partition a collection of data points into disjoint clusters. Often a bound $k$ on the number of clusters is given or assumed by the practitioner. Many…

机器学习 · 计算机科学 2012-02-01 Nir Ailon , Ron Begleiter

Several clustering methods (e.g., Normalized Cut and Ratio Cut) divide the Min Cut cost function by a cluster dependent factor (e.g., the size or the degree of the clusters), in order to yield a more balanced partitioning. We, instead,…

机器学习 · 计算机科学 2025-02-06 Morteza Haghir Chehreghani

The abundance of training data is not guaranteed in various supervised learning applications. One of these situations is the post-earthquake regional damage assessment of buildings. Querying the damage label of each building requires a…

机器学习 · 计算机科学 2021-08-17 Mohamadreza Sheibani , Ge Ou

We describe how to calculate standard errors for A/B tests that include clustered data, ratio metrics, and/or covariate adjustment. We may do this for power analysis/sample size calculations prior to running an experiment using historical…

统计方法学 · 统计学 2024-06-12 Tim Hesterberg , Ben Knight

We study constrained clustering, where constraints guide the clustering process. In existing works, two categories of constraints have been widely explored, namely pairwise and cardinality constraints. Pairwise constraints enforce the…

机器学习 · 计算机科学 2023-01-30 Adel Bibi , Ali Alqahtani , Bernard Ghanem

We introduce the Integrated Tsallis Combination (ITC), a hybrid impurity measure for decision tree learning that combines normalized Tsallis entropy with an exponential polarization component. While many existing measures sacrifice…

机器学习 · 统计学 2026-03-17 Edouard Lansiaux , Idriss Jairi , Hayfa Zgaya-Biau

There is a growing need for pluralistic alignment methods that can steer language models towards individual attributes and preferences. One such method, Self-Supervised Alignment with Mutual Information (SAMI), uses conditional mutual…

计算与语言 · 计算机科学 2025-06-06 Soham V. Govande

In an earlier study, we showed that Tsallis relative entropy (TRE), which is the generalization of Kullback-Leibler relative entropy (KLRE) to non-extensive systems, can be used as a possible risk measure in constructing risk optimal…

统计金融 · 定量金融 2022-05-30 Sandhya Devi , Sherman Page

The information theoretic quantity known as mutual information finds wide use in classification and community detection analyses to compare two classifications of the same set of objects into groups. In the context of classification…

社会与信息网络 · 计算机科学 2020-04-29 M. E. J. Newman , George T. Cantwell , Jean-Gabriel Young

This paper provides a user's guide to the general theory of approximate randomization tests developed in Canay, Romano, and Shaikh (2017) when specialized to linear regressions with clustered data. An important feature of the methodology is…

计量经济学 · 经济学 2022-03-16 Yong Cai , Ivan A. Canay , Deborah Kim , Azeem M. Shaikh

Clustering algorithms are used extensively in data analysis for data exploration and discovery. Technological advancements lead to continually growth of data in terms of volume, dimensionality and complexity. This provides great…

机器学习 · 计算机科学 2024-02-20 Miles McCrory , Spencer A. Thomas

We study methods for aggregating pairwise comparison data in order to estimate outcome probabilities for future comparisons among a collection of n items. Working within a flexible framework that imposes only a form of strong stochastic…

机器学习 · 计算机科学 2016-03-23 Nihar B. Shah , Sivaraman Balakrishnan , Martin J. Wainwright

In many health policy settings, adaptive interventions target a population of clusters (e.g., schools), with the ultimate intent of impacting outcomes at the level of individuals within the clusters. Health policy researchers can use…