中文
相关论文

相关论文: Assessing quality of selection procedures: Lower b…

200 篇论文

We introduce an Integrative Ranking and Thresholding (IRT) framework for fusing evidence from multiple testing procedures. The key innovation is a method that transforms binary testing decisions into compound $e-$values, enabling the…

统计方法学 · 统计学 2025-09-04 Trambak Banerjee , Bowen Gang , Jianliang He

In natural language processing (NLP) we always rely on human judgement as the golden quality evaluation method. However, there has been an ongoing debate on how to better evaluate inter-rater reliability (IRR) levels for certain evaluation…

计算与语言 · 计算机科学 2023-07-11 Serge Gladkoff , Lifeng Han , Goran Nenadic

The ongoing rapid development of the e-commercial and interest-base websites make it more pressing to evaluate objects' accurate quality before recommendation by employing an effective reputation system. The objects' quality are often…

物理与社会 · 物理学 2018-07-23 Leilei Wu , Zhuoming Ren , Xiao-Long Ren , Jianlin Zhang , Linyuan Lü

Binary classification is a task that involves the classification of data into one of two distinct classes. It is widely utilized in various fields. However, conventional classifiers tend to make overconfident predictions for data that…

机器学习 · 计算机科学 2025-03-13 Shoma Yokura , Akihisa Ichiki

In many real-world settings, the critical class is rare and a missed detection carries a disproportionately high cost. For example, tumors are rare and a false negative diagnosis could have severe consequences on treatment outcomes;…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Mohammadi Kiarash , Zhao He , Mengyao Zhai , Frederick Tung

Inverse Ising inference allows pairwise interactions of complex binary systems to be reconstructed from empirical correlations. Typical estimators used for this inference, such as Pseudo-likelihood maximization (PLM), are biased. Using the…

无序系统与神经网络 · 物理学 2023-07-19 Maximilian Benedikt Kloucek , Thomas Machon , Shogo Kajimura , C. Patrick Royall , Naoki Masuda , Francesco Turci

Machine learning models are increasingly used in critical decision-making applications. However, these models are susceptible to replicating or even amplifying bias present in real-world data. While there are various bias mitigation methods…

机器学习 · 计算机科学 2024-01-05 Shih-Chi Ma , Tatiana Ermakova , Benjamin Fabian

In this paper we address the problem of matching patterns in the so-called verification setting in which a novel, query pattern is verified against a single training pattern: the decision sought is whether the two match (i.e. belong to the…

计算机视觉与模式识别 · 计算机科学 2014-07-07 Ognjen Arandjelovic

Diagnostic tests play a crucial role in medical care. Thus any new diagnostic tests must undergo a thorough evaluation. New diagnostic tests are evaluated in comparison with the respective gold standard tests. The performance of binary…

应用统计 · 统计学 2025-09-17 Wan Nor Arifin , Umi Kalsom Yusof

We consider the problem of estimating the false-/ true-positive-rate (FPR/TPR) for a binary classification model when there are incorrect labels (label noise) in the validation set. Our motivating application is fraud prevention where…

机器学习 · 计算机科学 2023-08-08 Justin Tittelfitz

Relevance judgment of human assessors is inherently subjective and dynamic when evaluation datasets are created for Information Retrieval (IR) systems. However, a small group of experts' relevance judgment results are usually taken as…

信息检索 · 计算机科学 2022-08-09 Dengya Zhu , Shastri L Nimmagadda , Kok Wai Wong , Torsten Reiners

Reciprocal recommender systems~(RRS), conducting bilateral recommendations between two involved parties, have gained increasing attention for enhancing matching efficiency. However, the majority of existing methods in the literature still…

信息检索 · 计算机科学 2024-08-20 Chen Yang , Sunhao Dai , Yupeng Hou , Wayne Xin Zhao , Jun Xu , Yang Song , Hengshu Zhu

The Brier score is a widely used metric evaluating overall performance of probabilistic predictions for binary outcomes in clinical research. However, its interpretation can be complex, as it does not align with commonly taught concepts in…

应用统计 · 统计学 2025-07-08 Linard Hoessly

We present a new approach to interpreting IRR that is empirical and contextualized. It is based upon benchmarking IRR against baseline measures in a replication, one of which is a novel cross-replication reliability (xRR) measure based on…

应用统计 · 统计学 2021-06-15 Ka Wong , Praveen Paritosh , Lora Aroyo

Likelihood-to-evidence ratio estimation is usually cast as either a binary (NRE-A) or a multiclass (NRE-B) classification task. In contrast to the binary classification framework, the current formulation of the multiclass version has an…

机器学习 · 统计学 2024-07-08 Benjamin Kurt Miller , Christoph Weniger , Patrick Forré

Direct optimization of IR metrics has often been adopted as an approach to devise and develop ranking-based recommender systems. Most methods following this approach aim at optimizing the same metric being used for evaluation, under the…

信息检索 · 计算机科学 2021-06-07 Roger Zhe Li , Julián Urbano , Alan Hanjalic

The search engine evaluation research has quite a lot metrics available to it. Only recently, the question of the significance of individual metrics started being raised, as these metrics' correlations to real-world user experiences or…

信息检索 · 计算机科学 2013-02-12 Pavel Sirotkin

The principle of peer review is central to the evaluation of research, by ensuring that only high-quality items are funded or published. But peer review has also received criticism, as the selection of reviewers may introduce biases in the…

其他统计学 · 统计学 2015-07-24 Olivier Francois

Classification is a fundamental task in many applications on which data-driven methods have shown outstanding performances. However, it is challenging to determine whether such methods have achieved the optimal performance. This is mainly…

机器学习 · 计算机科学 2024-01-30 Minoh Jeong , Martina Cardone , Alex Dytso

Our study revisits the problem of accuracy-fairness tradeoff in binary classification. We argue that comparison of non-discriminatory classifiers needs to account for different rates of positive predictions, otherwise conclusions about…

机器学习 · 计算机科学 2015-05-22 Indre Zliobaite