中文
相关论文

相关论文: Concentration Inequalities for Two-Sample Rank Pro…

200 篇论文

Many classification performance metrics exist, each suited to a specific application. However, these metrics often differ in scale and can exhibit varying sensitivity to class imbalance rates in the test set. As a result, it is difficult to…

机器学习 · 统计学 2026-04-21 Ningsheng Zhao , Trang Bui , Jia Yuan Yu , Krzysztof Dzieciolowski

Traditional machine learning follows a close-set assumption that the training and test set share the same label space. While in many practical scenarios, it is inevitable that some test samples belong to unknown classes (open-set). To fix…

机器学习 · 计算机科学 2023-02-23 Zitai Wang , Qianqian Xu , Zhiyong Yang , Yuan He , Xiaochun Cao , Qingming Huang

Underspecification and fairness in machine learning (ML) applications have recently become two prominent issues in the ML community. Acoustic scene classification (ASC) applications have so far remained unaffected by this discussion, but…

机器学习 · 计算机科学 2021-10-05 Andreas Triantafyllopoulos , Manuel Milling , Konstantinos Drossos , Björn W. Schuller

We discuss a graph-based approach for testing spatial point patterns. This approach falls under the category of data-random graphs, which have been introduced and used for statistical pattern recognition in recent years. Our goal is to test…

统计方法学 · 统计学 2008-02-06 E. Ceyhan , C. E. Priebe , D. J. Marchette

So-called linear rank statistics provide a means for distribution-free (even in finite samples), yet highly flexible, two-sample testing in the setting of univariate random variables. Their flexibility derives from a choice of weights that…

统计方法学 · 统计学 2023-10-03 Dan D. Erdmann-Pham

The advent of modern data collection and processing techniques has seen the size, scale, and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the…

应用统计 · 统计学 2024-07-30 Zeyi Wang , Eric Bridgeford , Shangsi Wang , Joshua T. Vogelstein , Brian Caffo

The Kendall plot ($\K$-plot) is a plot measuring dependence between the components of a bivariate random variable. The $\K$-plot graphs the Kendall distribution function against the distribution function of $VU$, where $V$ and $U$ are…

统计理论 · 数学 2018-11-22 Albert Vexler , Georgios Afendras , Marianthi Markatou

Detecting and locating changes in highly multivariate data is a major concern in several current statistical applications. In this context, the first contribution of the paper is a novel non-parametric two-sample homogeneity test for…

统计理论 · 数学 2012-02-13 Alexandre Lung-Yut-Fong , Céline Lévy-Leduc , Olivier Cappé

Graph classification aims to categorize graphs based on their structural and attribute features, with applications in diverse fields such as social network analysis and bioinformatics. Among the methods proposed to solve this task, those…

机器学习 · 计算机科学 2025-07-23 Lucas Potin , Rosa Figueiredo , Vincent Labatut , Christine Largeron

The estimated accuracy of a classifier is a random quantity with variability. A common practice in supervised machine learning, is thus to test if the estimated accuracy is significantly better than chance level. This method of signal…

统计方法学 · 统计学 2020-01-28 Jonathan D. Rosenblatt , Yuval Benjamini , Roee Gilron , Roy Mukamel , Jelle J. Goeman

This paper investigates a statistical procedure for testing the equality of two independent estimated covariance matrices when the number of potentially dependent data vectors is large and proportional to the size of the vectors, that is,…

统计理论 · 数学 2020-06-01 Rémy Mariétan , Stephan Morgenthaler

Bipartite ranking is a fundamental ranking problem that learns to order relevant instances ahead of irrelevant ones. The pair-wise approach for bi-partite ranking construct a quadratic number of pairs to solve the problem, which is…

机器学习 · 计算机科学 2017-08-25 Wei-Yuan Shen , Hsuan-Tien Lin

In a typical Internet-of-Things setting that involves scientific applications, a target computation can be evaluated in many different ways depending on the split of computations among various devices. On the one hand, different…

性能 · 计算机科学 2022-08-09 Aravind Sankaran , Paolo Bientinesi

Real-world classification domains, such as medicine, health and safety, and finance, often exhibit imbalanced class priors and have asynchronous misclassification costs. In such cases, the classification model must achieve a high recall…

机器学习 · 计算机科学 2021-05-11 Michał Koziarski , Colin Bellinger , Michał Woźniak

The ROC curve is widely used to assess binary classifiers. Yet for some applications, such as alert systems for monitoring hospitalized patients, conventional ROC analysis cannot meet two key deployment needs: enforcing a constraint on…

机器学习 · 计算机科学 2026-04-03 Christopher Ratigan , Kyle Heuton , Carissa Wang , Lenore Cowen , Michael C. Hughes

Extending to dimension 2 and higher the dual univariate concepts of ranks and quantiles has remained an open problem for more than half a century. Based on measure transportation results, a solution has been proposed recently under the name…

统计理论 · 数学 2021-11-10 Marc Hallin , Gilles Mordant

Commonly used evaluation measures including Recall, Precision, F-Measure and Rand Accuracy are biased and should not be used without clear understanding of the biases, and corresponding identification of chance or base case levels of the…

机器学习 · 计算机科学 2020-11-02 David M. W. Powers

The ability to collect and store ever more massive databases has been accompanied by the need to process them efficiently. In many cases, most observations have the same behavior, while a probable small proportion of these observations are…

统计理论 · 数学 2021-09-21 Myrto Limnios , Nathan Noiry , Stéphan Clémençon

The Abstraction and Reasoning Corpus (ARC) is a visual program synthesis benchmark designed to test challenging out-of-distribution generalization in humans and machines. Since 2019, limited progress has been observed on the challenge using…

人工智能 · 计算机科学 2024-09-04 Solim LeGris , Wai Keen Vong , Brenden M. Lake , Todd M. Gureckis

While the area under the ROC curve is perhaps the most common measure that is used to rank the relative performance of different binary classifiers, longstanding field folklore has noted that it can be a measure that ill-captures the…

机器学习 · 计算机科学 2024-12-19 Christopher Ratigan , Lenore Cowen