中文
相关论文

相关论文: Quantifying Officiating Impact in the NBA: A Refer…

200 篇论文

Evaluating rare-event forecasts is challenging because standard metrics collapse as event prevalence declines. Measures such as F1-score, AUPRC, MCC, and accuracy induce degenerate thresholds -- converging to zero or one -- and their values…

统计方法学 · 统计学 2025-12-02 Sotirios D. Nikolopoulos

The purpose of this article is to develop a general parametric estimation theory that allows the derivation of the limit distribution of estimators in non-regular models where the true parameter value may lie on the boundary of the…

统计理论 · 数学 2022-11-28 Junichiro Yoshida , Nakahiro Yoshida

Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they are actually rewarding. In this paper we develop…

计算与语言 · 计算机科学 2025-05-21 David Reber , Sean Richardson , Todd Nief , Cristina Garbacea , Victor Veitch

Current IR evaluation is based on relevance judgments, created either manually or automatically, with decisions outsourced to Large Language Models (LLMs). We offer an alternative paradigm, that never relies on relevance judgments in any…

信息检索 · 计算机科学 2024-02-02 Naghmeh Farzi , Laura Dietz

Large language models are increasingly used to support high-stakes decisions, potentially influencing who is granted bail or receives a loan. Naive chain-of-thought sampling can improve average decision accuracy, but has also been shown to…

机器学习 · 计算机科学 2025-07-16 Zara Hall , Melanie Subbiah , Thomas P Zollo , Kathleen McKeown , Richard Zemel

Optimizing resistance training for hypertrophy requires balancing proximity to muscular failure, often quantified by Repetitions in Reserve (RiR), with fatigue management. However, subjective RiR assessment is unreliable, leading to…

机器学习 · 计算机科学 2025-12-16 Grant King , Musa Azeem , Savannah Noblitt , Ramtin Zand , Homayoun Valafar

In observational studies of discrimination, the most common statistical approaches consider either the rate at which decisions are made (benchmark tests) or the success rate of those decisions (outcome tests). Both tests, however, have…

应用统计 · 统计学 2025-03-07 Johann D. Gaebler , Sharad Goel

Recent advances in reinforcement learning (RL) have led to substantial improvements in the mathematical reasoning abilities of LLMs, as measured by standard benchmarks. Yet these gains often persist even when models are trained with flawed…

人工智能 · 计算机科学 2026-01-06 Jian Yao , Ran Cheng , Kay Chen Tan

Generating free-text rationales is a promising step towards explainable NLP, yet evaluating such rationales remains a challenge. Existing metrics have mostly focused on measuring the association between the rationale and a given label. We…

计算与语言 · 计算机科学 2023-06-05 Hanjie Chen , Faeze Brahman , Xiang Ren , Yangfeng Ji , Yejin Choi , Swabha Swayamdipta

Physics-informed statistical learning (PISL) integrates empirical data with physical knowledge to enhance the statistical performance of estimators. While PISL methods are widely used in practice, a comprehensive theoretical understanding…

机器学习 · 统计学 2025-10-28 Diego Marcondes

The Invariant Risk Minimization (IRM) framework aims to learn invariant features from a set of environments for solving the out-of-distribution (OOD) generalization problem. The underlying assumption is that the causal components of the…

机器学习 · 计算机科学 2021-12-28 Moulik Choraria , Ibtihal Ferwana , Ankur Mani , Lav R. Varshney

Reliability of machine learning evaluation -- the consistency of observed evaluation scores across replicated model training runs -- is affected by several sources of nondeterminism which can be regarded as measurement noise. Current…

机器学习 · 计算机科学 2023-10-10 Michael Hagmann , Philipp Meier , Stefan Riezler

Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-data performance of the full tune-then-deploy procedure. We…

机器学习 · 统计学 2026-05-08 Yang Xu , Jiefu Zhang , Haixiang Sun , Zihan Zhou , Tianyu Cao , Vaneet Aggarwal

Shooting skill in the NBA is typically measured by field goal percentage (FG%) - the number of makes out of the total number of shots. Even more advanced metrics like true shooting percentage are calculated by counting each player's…

应用统计 · 统计学 2018-10-30 Daniel Daly-Grafstein , Luke Bornn

Comprehensive evaluations of language models (LM) during both development and deployment phases are necessary because these models possess numerous capabilities (e.g., mathematical reasoning, legal support, or medical diagnostic) as well as…

计算与语言 · 计算机科学 2025-03-18 Sang Truong , Yuheng Tu , Percy Liang , Bo Li , Sanmi Koyejo

Formula 1 performance is a combination of the car's ability and the driver's ability. While a given race or season can tell you how well a car and driver performed jointly, isolating the individual impact of the driver and constructor…

应用统计 · 统计学 2025-08-04 Saurabh Rane

Large language models are often used as judges to score candidate responses, then validated with a single global metric such as correlation with reference labels. This can be misleading when the real deployment task is best-of-n selection…

机器学习 · 计算机科学 2026-03-16 Eddie Landesberg

We propose a restricted win probability estimand for comparing treatments in a randomized trial with a time-to-event outcome. We also propose Bayesian estimators for this summary measure as well as the unrestricted win probability. Bayesian…

统计方法学 · 统计学 2024-11-06 Michelle Leeberg , Xianghua Luo , Thomas A. Murray

Reinforcement learning (RL) systems have countless applications, from energy-grid management to protein design. However, such real-world scenarios are often extremely difficult, combinatorial in nature, and require complex coordination…

Defensive Pass Interference (DPI) is one of the most impactful penalties in the NFL. DPI is a spot foul, yielding an automatic first down to the team in possession. With such an influence on the game, referees have no room for a mistake. It…

机器学习 · 计算机科学 2022-06-28 Arian Skoki , Jonatan Lerga , Ivan Štajduhar