English
Related papers

Related papers: Position: Stop Chasing the C-index when Evaluating…

200 papers

Pairwise model comparisons drawn from foundation-model benchmarks ("A is safer than B") are read as quantitative verdicts but hinge on harness choices benchmark papers under-specify. We close one theory-benchmark loop on this primitive: a…

Machine Learning · Computer Science 2026-05-26 Yanhang Li , Zhichao Fan , Zexin Zhuang

Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the layer where alignment often operates: routing from concept detection to behavioral policy. We…

Machine Learning · Computer Science 2026-05-04 Gregory N. Frank

Benchmarking models is a key factor for the rapid progress in machine learning (ML) research. Thus, further progress depends on improving benchmarking metrics. A standard metric to measure the behavioral alignment between ML models and…

Neurons and Cognition · Quantitative Biology 2025-11-10 Thomas Klein , Sascha Meyen , Wieland Brendel , Felix A. Wichmann , Kristof Meding

Evaluating alignment in language models requires testing how they behave under realistic pressure, not just what they claim they would do. While alignment failures increasingly cause real-world harm, comprehensive evaluation frameworks with…

Artificial Intelligence · Computer Science 2026-02-25 Nora Petrova , John Burden

Survival models capture the relationship between an accumulating hazard and the occurrence of a singular event stimulated by that accumulation. When the model for the hazard is sufficiently flexible survival models can accommodate a wide…

Methodology · Statistics 2024-03-04 Michael Betancourt

In recent machine learning systems, confidence scores are being utilized more and more to manage selective prediction, whereby a model can abstain from making a prediction when it is unconfident. Yet, conventional metrics like accuracy,…

Machine Learning · Computer Science 2025-05-27 Kourosh Shahnazari , Seyed Moein Ayyoubzadeh , Mohammadali Keshtparvar , Pegah Ghaffari

Under a single-index regression assumption, we introduce a new semiparametric procedure to estimate a conditional density of a censored response. The regression model can be seen as a generalization of Cox regression model and also as a…

Statistics Theory · Mathematics 2009-03-22 Olivier Bouaziz , Olivier Lopez

The evaluation of supervised machine learning models is a critical stage in the development of reliable predictive systems. Despite the widespread availability of machine learning libraries and automated workflows, model assessment is often…

Machine Learning · Computer Science 2026-04-16 Xuanyan Liu , Ignacio Cabrera Martin , Marcello Trovati , Xiaolong Xu , Nikolaos Polatidis

Modeling symptom progression to identify informative subjects for a new Huntington's disease clinical trial is problematic since time to diagnosis, a key covariate, can be heavily censored. Imputation is an appealing strategy where censored…

Methodology · Statistics 2025-02-11 Sarah C. Lotspeich , Tanya P. Garcia

Factorial analyses offer a powerful nonparametric means to detect main or interaction effects among multiple treatments. For survival outcomes, e.g. from clinical trials, such techniques can be adopted for comparing reasonable…

Methodology · Statistics 2023-02-06 Takeshi Emura , Marc Ditzhaus , Dennis Dobler , Kenta Murotani

We analyse an issue when comparing survival curves between two subgroups. We show that there is a direct relationship between estimates of subgroups' survival at a time point and positive and negative predictive values in the binary…

Methodology · Statistics 2021-05-17 Damjan Krstajic

The search engine evaluation research has quite a lot metrics available to it. Only recently, the question of the significance of individual metrics started being raised, as these metrics' correlations to real-world user experiences or…

Information Retrieval · Computer Science 2013-02-12 Pavel Sirotkin

Applications of machine learning in healthcare often require working with time-to-event prediction tasks including prognostication of an adverse event, re-hospitalization or death. Such outcomes are typically subject to censoring due to…

Machine Learning · Computer Science 2022-08-04 Chirag Nagpal , Willa Potosnak , Artur Dubrawski

Ranked set sampling (RSS) is a cost-efficient study design that uses inexpensive baseline ranking to select a more informative subset of individuals for full measurement. While RSS is well known to improve precision over simple random…

Methodology · Statistics 2025-12-30 Nabil Awan , Richard J. Chappell

Word and sentence embeddings are useful feature representations in natural language processing. However, intrinsic evaluation for embeddings lags far behind, and there has been no significant update since the past decade. Word and sentence…

Computation and Language · Computer Science 2022-03-22 Bin Wang , C. -C. Jay Kuo , Haizhou Li

This position paper argues that the evaluation of modern visual processing systems should no longer be driven primarily by single-metric image quality assessment benchmarks, particularly in the era of generative and perception-oriented…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Jinfan Hu , Fanghua Yu , Zhiyuan You , Xiang Yin , Hongyu An , Xinqi Lin , Chao Dong , Jinjin Gu

Uncertainty quantification of prediction models through prediction sets is increasingly popular and successful, but most existing methods rely on directly observing the outcome and do not appropriately handle censored outcomes, such as…

Methodology · Statistics 2025-05-06 Wenwen Si , Hongxiang Qiu

Instruction following aims to align Large Language Models (LLMs) with human intent by specifying explicit constraints on how tasks should be performed. However, we reveal a counterintuitive phenomenon: instruction following can…

Computation and Language · Computer Science 2026-01-30 Yunjia Qi , Hao Peng , Xintong Shi , Amy Xin , Xiaozhi Wang , Bin Xu , Lei Hou , Juanzi Li

We present a methodology for model evaluation and selection where the sampling mechanism violates the i.i.d. assumption. Our methodology involves a formulation of the bias between the standard Cross-Validation (CV) estimator and the mean…

Methodology · Statistics 2025-03-14 Oren Yuval , Saharon Rosset

Reward-model-based fine-tuning is a central paradigm in aligning Large Language Models with human preferences. However, such approaches critically rely on the assumption that proxy reward models accurately reflect intended supervision, a…

Computation and Language · Computer Science 2026-01-21 Zixuan Liu , Siavash H. Khajavi , Guangkai Jiang , Xinru Liu