中文
相关论文

相关论文: The Divergence Index: A Decomposable Measure of Se…

200 篇论文

Disparities in health or well-being experienced by minority groups can be difficult to study using the traditional exposure-outcome paradigm in causal inference, since potential outcomes in variables such as race or sexual minority status…

统计方法学 · 统计学 2025-01-22 Andy A. Shen , Elina Visoki , Ran Barzilay , Samuel D. Pimentel

Partial Information Decomposition (PID) has become one of the most prominent information-theoretic frameworks for describing the structure and quality of information in complex systems. Despite its widespread utility, there exists no unique…

信息论 · 计算机科学 2026-03-10 Alberto Liardi , Keenan J. A. Down , George Blackburne , Matteo Neri , Pedro A. M. Mediano

We introduce a cluster evaluation technique called Tree Index. Our Tree Index algorithm aims at describing the structural information of the clustering rather than the quantitative format of cluster-quality indexes (where the representation…

机器学习 · 计算机科学 2020-03-25 A. H. Beg , Md Zahidul Islam , Vladimir Estivill-Castro

We study a separable design for computing information measures, where the information measure is computed from learned feature representations instead of raw data. Under mild assumptions on the feature representations, we demonstrate that a…

信息论 · 计算机科学 2025-01-28 Xiangxiang Xu , Lizhong Zheng

A system's heterogeneity (\textit{diversity}) is the effective size of its event space, and can be quantified using the R\'enyi family of indices (also known as Hill numbers in ecology or Hannah-Kay indices in economics), which are indexed…

定量方法 · 定量生物学 2020-08-26 Abraham Nunes , Martin Alda , Thomas Trappenberg

Size uniformity is one of the main criteria of superpixel methods. But size uniformity rarely conforms to the varying content of an image. The chosen size of the superpixels therefore represents a compromise - how to obtain the fewest…

计算机视觉与模式识别 · 计算机科学 2016-11-29 Radhakrishna Achanta , Pablo Márquez-Neila , Pascal Fua , Sabine Süsstrunk

A significant body of research in the data sciences considers unfair discrimination against social categories such as race or gender that could occur or be amplified as a result of algorithmic decisions. Simultaneously, real-world…

机器学习 · 计算机科学 2022-12-09 Lucius E. J. Bynum , Joshua R. Loftus , Julia Stoyanovich

While Kolmogorov complexity is the accepted absolute measure of information content in an individual finite object, a similarly absolute notion is needed for the information distance between two individual objects, for example, two…

信息论 · 计算机科学 2010-06-18 Charles H. Bennett , Peter Gacs , Ming Li , Paul M. B. Vitanyi , Wojciech H. Zurek

In search and recommendation, diversifying the multi-aspect search results could help with reducing redundancy, and promoting results that might not be shown otherwise. Many previous methods have been proposed for this task. However,…

信息检索 · 计算机科学 2021-05-24 Jianghong Zhou , Eugene Agichtein , Surya Kallumadi

Algorithmic fairness is becoming increasingly important in data mining and machine learning. Among others, a foundational notation is group fairness. The vast majority of the existing works on group fairness, with a few exceptions,…

机器学习 · 计算机科学 2023-01-03 Jian Kang , Tiankai Xie , Xintao Wu , Ross Maciejewski , Hanghang Tong

Information collection is a fundamental problem in big data, where the size of sampling sets plays a very important role. This work considers the information collection process by taking message importance into account. Similar to…

信息论 · 计算机科学 2018-01-15 Shanyun Liu , Rui She , Pingyi Fan

The paper considers a new quantitative-qualitative proximity measure for the features of information objects, where data enters a common information resource from several sources independently. The goal is to determine the possibility of…

人工智能 · 计算机科学 2026-04-08 Volodymyr Yuzefovych

In distributed and federated learning, heterogeneity across data sources remains a major obstacle to effective model aggregation and convergence. We focus on feature heterogeneity and introduce energy distance as a sensitive measure for…

机器学习 · 统计学 2025-01-28 Mengchen Fan , Baocheng Geng , Roman Shterenberg , Joseph A. Casey , Zhong Chen , Keren Li

Urban segregation refers to the physical and social division of people, often driving inequalities within cities and exacerbating socioeconomic and racial tensions. While most studies focus on residential spaces, they often neglect…

人机交互 · 计算机科学 2025-01-08 Yue Yu , Yifang Wang , Yongjun Zhang , Huamin Qu , Dongyu Liu

Clustering has been a major research topic in the field of machine learning, one to which Deep Learning has recently been applied with significant success. However, an aspect of clustering that is not addressed by existing deep clustering…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Ioannis Maniadis Metaxas , Georgios Tzimiropoulos , Ioannis Patras

This paper is an attempt to set a justification for making use of some dicrepancy indexes, starting from the classical Maximum Likelihood definition, and adapting the corresponding basic principle of inference to situations where…

统计理论 · 数学 2021-02-24 Michel Broniatowski

A new interpoint distance-based measure is proposed to identify the optimal number of clusters present in a data set. Designed in nonparametric approach, it is independent of the distribution of given data. Interpoint distances between the…

机器学习 · 计算机科学 2022-10-18 Soumita Modak

Intuitively, the concept of similarity is the notion to measure an inexact matching between two entities of the same reference set. The notions of similarity and its close relative dissimilarity are widely used in many fields of Artificial…

人工智能 · 计算机科学 2012-12-13 Lluís A. Belanche

In the context of machine learning, disparate impact refers to a form of systematic discrimination whereby the output distribution of a model depends on the value of a sensitive attribute (e.g., race or gender). In this paper, we propose an…

信息论 · 计算机科学 2018-05-14 Hao Wang , Berk Ustun , Flavio P. Calmon

In machine learning, observation features are measured in a metric space to obtain their distance function for optimization. Given similar features that are statistically sufficient as a population, a statistical distance between two…

机器学习 · 统计学 2020-06-23 Xin Lu