中文
相关论文

相关论文: Ground Truth Bias in External Cluster Validity Ind…

200 篇论文

Confidence intervals (CIs) are instrumental in statistical analysis, providing a range estimate of the parameters. In modern statistics, selective inference is common, where only certain parameters are highlighted. However, this selective…

统计方法学 · 统计学 2025-09-17 Tzviel Frostig , Yoav Benjamini , Ruth Heller

Data clustering involves identifying latent similarities within a dataset and organizing them into clusters or groups. The outcomes of various clustering algorithms differ as they are susceptible to the intrinsic characteristics of the…

机器学习 · 计算机科学 2024-07-31 Bryar A. Hassan , Noor Bahjat Tayfor , Alla A. Hassan , Aram M. Ahmed , Tarik A. Rashid , Naz N. Abdalla

A growing literature on human-AI decision-making investigates strategies for combining human judgment with statistical models to improve decision-making. Research in this area often evaluates proposed improvements to models, interfaces, or…

计算机与社会 · 计算机科学 2023-05-29 Luke Guerdan , Amanda Coston , Zhiwei Steven Wu , Kenneth Holstein

We consider the simultaneous clustering of rows and columns of a matrix and more particularly the ability to measure the agreement between two co-clustering partitions. The new criterion we developed is based on the Adjusted Rand Index and…

应用统计 · 统计学 2020-12-16 Valerie Robert , Yann Vasseur , Vincent Brault

Stochastic Natural Gradient Variational Inference (NGVI) is a widely used method for approximating posterior distribution in probabilistic models. Despite its empirical success and foundational role in variational inference, its theoretical…

机器学习 · 计算机科学 2025-10-23 Fangyuan Sun , Ilyas Fatkhullin , Niao He

Shift invariance is a critical property of CNNs that improves performance on classification. However, we show that invariance to circular shifts can also lead to greater sensitivity to adversarial attacks. We first characterize the margin…

机器学习 · 计算机科学 2021-11-23 Songwei Ge , Vasu Singla , Ronen Basri , David Jacobs

Gradient Boosting Machines (GBM) are among the go-to algorithms on tabular data, which produce state of the art results in many prediction tasks. Despite its popularity, the GBM framework suffers from a fundamental flaw in its base…

机器学习 · 计算机科学 2021-09-14 Afek Ilay Adler , Amichai Painsky

As a computational alternative to Markov chain Monte Carlo approaches, variational inference (VI) is becoming more and more popular for approximating intractable posterior distributions in large-scale Bayesian models due to its comparable…

机器学习 · 统计学 2023-06-05 Anirban Bhattacharya , Debdeep Pati , Yun Yang

Cross-Domain Recommendation (CDR) aims to leverage knowledge from a relatively data-richer source domain to address the data sparsity problem in a relatively data-sparser target domain. While CDR methods need to address the distribution…

信息检索 · 计算机科学 2025-05-23 Jiajie Zhu , Yan Wang , Feng Zhu , Pengfei Ding , Hongyang Liu , Zhu Sun

Variable clustering is important for explanatory analysis. However, only few dedicated methods for variable clustering with the Gaussian graphical model have been proposed. Even more severe, small insignificant partial correlations due to…

应用统计 · 统计学 2018-06-18 Daniel Andrade , Akiko Takeda , Kenji Fukumizu

Internal cluster validity measures (such as the Calinski-Harabasz, Dunn, or Davies-Bouldin indices) are frequently used for selecting the appropriate number of partitions a dataset should be split into. In this paper we consider what…

机器学习 · 统计学 2022-08-31 Marek Gagolewski , Maciej Bartoszuk , Anna Cena

K-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very process of CV requires random partitioning of the data and so…

计算与语言 · 计算机科学 2018-06-20 Henry B. Moss , David S. Leslie , Paul Rayson

Machine learning models in dynamic environments often suffer from concept drift, where changes in the data distribution degrade performance. While detecting this drift is a well-studied topic, explaining how and why the model's…

机器学习 · 计算机科学 2025-09-12 Ignacy Stępka , Jerzy Stefanowski

Gait recognition aims to identify individuals by recognizing their walking patterns. However, an observation is made that most of the previous gait recognition methods degenerate significantly due to two memorization effects, namely…

计算机视觉与模式识别 · 计算机科学 2022-10-14 Weichen Yu , Hongyuan Yu , Yan Huang , Chunshui Cao , Liang Wang

Machine learning algorithms in socially sensitive domains (e.g., credit decisions) often focus on equalizing predictive outcomes. However, satisfying these metrics does not guarantee that models use the same reasoning for different groups.…

机器学习 · 计算机科学 2026-05-14 Gideon Popoola , John Sheppard

Determining the appropriate rank in Non-negative Matrix Factorization (NMF) is a critical challenge that often requires extensive parameter tuning and domain-specific knowledge. Traditional methods for rank determination focus on…

机器学习 · 计算机科学 2024-10-22 Marc A. Tunnell , Zachary J. DeBruine , Erin Carrier

Unsupervised person re-identification (ReID) aims to train a feature extractor for identity retrieval without exploiting identity labels. Due to the blind trust in imperfect clustering results, the learning is inevitably misled by…

计算机视觉与模式识别 · 计算机科学 2022-11-23 Yunqi Miao , Jiankang Deng , Guiguang Ding , Jungong Han

The integration of deep learning tools in gastrointestinal vision holds the potential for significant advancements in diagnosis, treatment, and overall patient care. A major challenge, however, is these tools' tendency to make overconfident…

We establish the asymptotic implicit bias of gradient descent (GD) for generic non-homogeneous deep networks under exponential loss. Specifically, we characterize three key properties of GD iterates starting from a sufficiently small…

机器学习 · 计算机科学 2025-07-17 Yuhang Cai , Kangjie Zhou , Jingfeng Wu , Song Mei , Michael Lindsey , Peter L. Bartlett

In this paper we revisit the bias-variance decomposition of model error from the perspective of designing a fair classifier: we are motivated by the widely held socio-technical belief that noise variance in large datasets in social domains…

机器学习 · 计算机科学 2023-02-20 Falaah Arif Khan , Julia Stoyanovich