中文
相关论文

相关论文: Concentration Inequalities for Two-Sample Rank Pro…

200 篇论文

Ranking and scoring are ubiquitous. We consider the setting in which an institution, called a ranker, evaluates a set of individuals based on demographic, behavioral or other characteristics. The final output is a ranking that represents…

数据库 · 计算机科学 2016-10-28 Ke Yang , Julia Stoyanovich

Multiclass classifiers are often designed and evaluated only on a sample from the classes on which they will eventually be applied. Hence, their final accuracy remains unknown. In this work we study how a classifier's performance over the…

机器学习 · 计算机科学 2024-05-29 Yuli Slavutsky , Yuval Benjamini

The runs test is a well-known test that is used for checking independence between elements of a sample data sequence. Some of runs tests are based on the longest run and others based on the total runs. In this paper, we consider order…

统计方法学 · 统计学 2014-10-31 Mohammad Reza Kazemi , Ali Akbar Jafari

Binary decisions are very common in artificial intelligence. Applying a threshold on the continuous score gives the human decider the power to control the operating point to separate the two classes. The classifier,s discriminating power is…

人工智能 · 计算机科学 2016-06-03 Paulo J. L. Adeodato , Sílvio B. Melo

Everybody writes that ROC curves, a very common tool in binary classification problems, should be optimal, and in particular concave, non-decreasing and above the 45-degree line. Everybody uses ROC curves, theoretical and especially…

统计方法学 · 统计学 2019-08-01 Lidia Sacchetto , Mauro Gasparini

In recommendation systems, one is interested in the ranking of the predicted items as opposed to other losses such as the mean squared error. Although a variety of ways to evaluate rankings exist in the literature, here we focus on the Area…

机器学习 · 统计学 2015-08-26 Charanpal Dhanjal , Romaric Gaudel , Stephan Clemencon

In observational studies of discrimination, the most common statistical approaches consider either the rate at which decisions are made (benchmark tests) or the success rate of those decisions (outcome tests). Both tests, however, have…

应用统计 · 统计学 2025-03-07 Johann D. Gaebler , Sharad Goel

Assessing the performance of a learned model is a crucial part of machine learning. However, in some domains only positive and unlabeled examples are available, which prohibits the use of most standard evaluation metrics. We propose an…

机器学习 · 统计学 2015-12-31 Marc Claesen , Jesse Davis , Frank De Smet , Bart De Moor

Most binary classifiers work by processing the input to produce a scalar response and comparing it to a threshold value. The various measures of classifier performance assume, explicitly or implicitly, probability distributions $P_s$ and…

机器学习 · 计算机科学 2019-09-24 Luma Omar , Ioannis Ivrissimtzis

We study a game theoretic model of standardized testing for college admissions. Students are of two types; High and Low. There is a college that would like to admit the High type students. Students take a potentially costly standardized…

计算机科学与博弈论 · 计算机科学 2021-02-17 Sampath Kannan , Mingzi Niu , Aaron Roth , Rakesh Vohra

The Area Under the ROC Curve (AUC) is a crucial metric for machine learning, which evaluates the average performance over all possible True Positive Rates (TPRs) and False Positive Rates (FPRs). Based on the knowledge that a skillful…

机器学习 · 计算机科学 2022-06-24 Zhiyong Yang , Qianqian Xu , Shilong Bao , Yuan He , Xiaochun Cao , Qingming Huang

Prior to clinical applications, it is critical that risk prediction models are evaluated in independent studies that did not contribute to model development. While prospective cohort studies provide a natural setting for model validation,…

统计方法学 · 统计学 2017-10-13 Parichoy Pal Choudhury , Anil K. Chaturvedi , Nilanjan Chatterjee

Robust classification algorithms have been developed in recent years with great success. We take advantage of this development and recast the classical two-sample test problem in the framework of classification. Based on the estimates of…

统计理论 · 数学 2019-09-18 Haiyan Cai , Bryan Goggin , Qingtang Jiang

We propose new simultaneous inference methods for diagnostic trials with elaborate factorial designs. Instead of the commonly used total area under the receiver operating characteristic (ROC) curve, our parameters of interest are partial…

统计理论 · 数学 2023-02-22 Maximilian Wechsung , Frank Konietschke

Unsupervised ranking faces one critical challenge in evaluation applications, that is, no ground truth is available. When PageRank and its variants show a good solution in related subjects, they are applicable only for ranking from…

机器学习 · 计算机科学 2014-02-20 Chun-Guo Li , Xing Mei , Bao-Gang Hu

Randomized Controlled Trials (RCTs) are the gold standard for comparing the effectiveness of a new treatment to the current one (the control). Most RCTs allocate the patients to the treatment group and the control group by uniform…

机器学习 · 统计学 2018-10-22 Onur Atan , William R. Zame , Mihaela van der Schaar

The one-sample log-rank test is the method of choice for single-arm Phase II trials with time-to-event endpoint. It allows to compare the survival of the patients to a reference survival curve that typically represents the expected survival…

统计方法学 · 统计学 2026-03-02 Jannik Feld , Moritz Fabian Danzer , Andreas Faldum , Rene Schmidt

Two-way partial AUC (TPAUC) is a critical performance metric for binary classification with imbalanced data, as it focuses on specific ranges of the true positive rate (TPR) and false positive rate (FPR). However, stochastic algorithms for…

机器学习 · 计算机科学 2025-09-30 Linli Zhou , Bokun Wang , My T. Thai , Tianbao Yang

The Receiver Operating Characteristic (ROC) curve stands as a cornerstone in assessing the efficacy of biomarkers for disease diagnosis. Beyond merely evaluating performance, it provides with an optimal cutoff for biomarker values, crucial…

统计方法学 · 统计学 2025-04-29 Soutik Ghosal

The comparison of alternative rankings of a set of items is a general and prominent task in applied statistics. Predictor variables are ranked according to magnitude of association with an outcome, prediction models rank subjects according…