中文
相关论文

相关论文: Over-optimism in benchmark studies and the multipl…

200 篇论文

Observational studies provide invaluable opportunities to draw causal inference, but they may suffer from biases due to pretreatment difference between treated and control units. Matching is a popular approach to reduce observed covariate…

统计方法学 · 统计学 2025-09-17 Xinran Li

Clinical trials often evaluate multiple outcome variables to form a comprehensive picture of the effects of a new treatment. The resulting multidimensional insight contributes to clinically relevant and efficient decision-making about…

统计方法学 · 统计学 2023-08-14 X. M. Kavelaars , J. Mulder , M. C. Kaptein

This paper presents a novel approach to analyze human decision-making that involves comparing the behavior of professional chess players relative to a computational benchmark of cognitively bounded rationality. This benchmark is constructed…

综合经济学 · 经济学 2020-12-03 Dainis Zegners , Uwe Sunde , Anthony Strittmatter

The last decades saw dramatic progress in brain research. These advances were often buttressed by probing single variables to make circumscribed discoveries, typically through null hypothesis significance testing. New ways for generating…

神经元与认知 · 定量生物学 2019-03-26 Danilo Bzdok , John Ioannidis

Optimisation algorithms are commonly compared on benchmarks to get insight into performance differences. However, it is not clear how closely benchmarks match the properties of real-world problems because these properties are largely…

神经与进化计算 · 计算机科学 2021-07-15 Koen van der Blom , Timo M. Deist , Vanessa Volz , Mariapia Marchi , Yusuke Nojima , Boris Naujoks , Akira Oyama , Tea Tušar

This paper offers a commentary on the use of notions of statistical significance in choice modelling. We review the reasons for uncertainty in parameter estimates, provide a precise discussion on the computation of measures of uncertainty…

计量经济学 · 经济学 2026-05-18 Stephane Hess , Andrew Daly , Michiel Bliemer , Angelo Guevara , Ricardo Daziano , Thijs Dekker

The constant development of new data analysis methods in many fields of research is accompanied by an increasing awareness that these new methods often perform better in their introductory paper than in subsequent comparison studies…

统计方法学 · 统计学 2024-01-17 Christina Nießl , Sabine Hoffmann , Theresa Ullmann , Anne-Laure Boulesteix

Predictive algorithms inform consequential decisions in settings with selective labels: outcomes are observed only for units selected by past decision makers. This creates an identification problem under unobserved confounding -- when…

计量经济学 · 经济学 2025-11-07 Ashesh Rambachan , Amanda Coston , Edward Kennedy

Methods for building fair predictors often involve tradeoffs between fairness and accuracy and between different fairness criteria, but the nature of these tradeoffs varies. Recent work seeks to characterize these tradeoffs in specific…

机器学习 · 统计学 2021-09-02 Alan Mishler , Edward Kennedy

Novel reinforcement learning algorithms, or improvements on existing ones, are commonly justified by evaluating their performance on benchmark environments and are compared to an ever-changing set of standard algorithms. However, despite…

机器学习 · 计算机科学 2024-06-25 Scott M. Jordan , Adam White , Bruno Castro da Silva , Martha White , Philip S. Thomas

In robust optimization, the uncertainty set is used to model all possible outcomes of uncertain parameters. In the classic setting, one assumes that this set is provided by the decision maker based on the data available to her. Only…

最优化与控制 · 数学 2019-01-23 Trivikram Dokka , Marc Goerigk , Rahul Roy

Causal inference with observational studies often suffers from unmeasured confounding, yielding biased estimators based on the unconfoundedness assumption. Sensitivity analysis assesses how the causal conclusions change with respect to…

统计方法学 · 统计学 2024-04-01 Sizhu Lu , Peng Ding

Deep learning models have proven to be highly successful. Yet, their over-parameterization gives rise to model multiplicity, a phenomenon in which multiple models achieve similar performance but exhibit distinct underlying behaviours. This…

机器学习 · 计算机科学 2023-11-28 Prakhar Ganesh

The work applies the funnel plot methodology to measure and visualize uncertainty in the research performance of Italian universities in the science disciplines. The performance assessment is carried out at both discipline and overall…

数字图书馆 · 计算机科学 2018-10-31 Giovanni Abramo , Ciriaco Andrea D'Angelo , Leonardo Grilli

Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark studies attempt to quantify uncertainty, partly due to the…

计算与语言 · 计算机科学 2025-06-30 Robert E. Blackwell , Jon Barry , Anthony G. Cohn

Insufficient performance of optimization approaches for fitting of mathematical models is still a major bottleneck in systems biology. In this manuscript, the reasons and methodological challenges are summarized as well as their impact in…

性能 · 计算机科学 2019-07-09 Clemens Kreutz

Machine learning (ML) is increasingly used in high-stakes settings, yet multiplicity - the existence of multiple good models - means that some predictions are essentially arbitrary. ML researchers and philosophers posit that multiplicity…

计算机与社会 · 计算机科学 2025-01-24 Anna P. Meyer , Yea-Seul Kim , Aws Albarghouthi , Loris D'Antoni

Statistical inferential results generally come with a measure of reliability for decision-making purposes. For a policy implementer, the value of implementing published policy research depends critically upon this reliability. For a policy…

其他统计学 · 统计学 2024-08-21 Duncan Ermini Leaf

We provide an approach to exploratory data analysis in matched observational studies with a single intervention and multiple endpoints. In such settings, the researcher would like to explore evidence for actual treatment effects among these…

统计方法学 · 统计学 2025-12-10 Mengqi Lin , Colin Fogarty

We examine multi-task benchmarks in machine learning through the lens of social choice theory. We draw an analogy between benchmarks and electoral systems, where models are candidates and tasks are voters. This suggests a distinction…

机器学习 · 计算机科学 2024-05-07 Guanhua Zhang , Moritz Hardt