中文
相关论文

相关论文: Blinded sample size re-estimation in equivalence t…

200 篇论文

We analyze theoretical properties of the hybrid test for superior predictability. We demonstrate with a simple example that the test may not be pointwise asymptotically of level $\alpha$ at commonly used significance levels and may lead to…

计量经济学 · 经济学 2021-09-13 Deborah Kim

Much research on Machine Learning testing relies on empirical studies that evaluate and show their potential. However, in this context empirical results are sensitive to a number of parameters that can adversely impact the results of the…

软件工程 · 计算机科学 2023-09-12 Salah Ghamizi , Maxime Cordy , Yuejun Guo , Mike Papadakis , And Yves Le Traon

Score-based divergences have been widely used in machine learning and statistics applications. Despite their empirical success, a blindness problem has been observed when using these for multi-modal distributions. In this work, we discuss…

机器学习 · 统计学 2025-11-25 Mingtian Zhang , Oscar Key , Peter Hayes , David Barber , Brooks Paige , François-Xavier Briol

Binary observations are often repeated to improve data quality, creating technical replicates. Several scoring methods are commonly used to infer the actual individual state and obtain a probability for each state. The common practice of…

统计方法学 · 统计学 2025-01-24 Manuela Royer-Carenzi , Hadrien Lorenzo , Pierre Pudlo

An ongoing "reproducibility crisis" calls into question scientific discoveries across a variety of disciplines ranging from life to social sciences. Replication studies aim to investigate the validity of findings in published research, and…

应用统计 · 统计学 2023-05-09 Konstantinos Bourazas , Guido Consonni , Laura Deldossi

We examine the problem of construction of confidence intervals within the basic single-parameter, single-iteration variation of the method of quasi-optimal weights. Two kinds of distortions of such intervals due to insufficiently large…

数据分析、统计与概率 · 物理学 2020-05-27 A. D. Morozov , A. V. Lokhov , F. V. Tkachov

The growing awareness of safety concerns in large language models (LLMs) has sparked considerable interest in the evaluation of safety. This study investigates an under-explored issue about the evaluation of LLMs, namely the substantial…

计算与语言 · 计算机科学 2024-04-02 Yixu Wang , Yan Teng , Kexin Huang , Chengqi Lyu , Songyang Zhang , Wenwei Zhang , Xingjun Ma , Yu-Gang Jiang , Yu Qiao , Yingchun Wang

Clustering is part of unsupervised analysis methods that consist in grouping samples into homogeneous and separate subgroups of observations also called clusters. To interpret the clusters, statistical hypothesis testing is often used to…

统计方法学 · 统计学 2022-10-25 Benjamin Hivert , Denis Agniel , Rodolphe Thiébaut , Boris P Hejblum

We introduce equivalence testing procedures for linear regression analyses. Such tests can be very useful for confirming the lack of a meaningful association between a continuous outcome and a continuous or binary predictor. Specifically,…

统计方法学 · 统计学 2023-05-17 Harlan Campbell

This paper is concerned with the simultaneous estimation of $k$ population means when one suspects that the $k$ means are nearly equal. As an alternative to the preliminary test estimator based on the test statistics for testing hypothesis…

统计理论 · 数学 2018-09-13 Ryo Imai , Tatsuya Kubokawa , Malay Ghosh

Cross-validation is a widely-used technique to estimate prediction error, but its behavior is complex and not fully understood. Ideally, one would like to think that cross-validation estimates the prediction error for the model at hand, fit…

统计方法学 · 统计学 2024-03-12 Stephen Bates , Trevor Hastie , Robert Tibshirani

Annotated datasets are an essential ingredient to train, evaluate, compare and productionalize supervised machine learning models. It is therefore imperative that annotations are of high quality. For their creation, good quality management…

机器学习 · 计算机科学 2024-05-30 Jan-Christoph Klie , Juan Haladjian , Marc Kirchner , Rahul Nair

The problem of binary hypothesis testing between two probability measures is considered. New sharp bounds are derived for the best achievable error probability of such tests based on independent and identically distributed observations.…

信息论 · 计算机科学 2024-05-30 Valentinian Lungu , Ioannis Kontoyiannis

Label noise in data has long been an important problem in supervised learning applications as it affects the effectiveness of many widely used classification methods. Recently, important real-world applications, such as medical diagnosis…

机器学习 · 统计学 2021-12-02 Shunan Yao , Bradley Rava , Xin Tong , Gareth James

A common issue for classification in scientific research and industry is the existence of imbalanced classes. When sample sizes of different classes are imbalanced in training data, naively implementing a classification method often leads…

统计方法学 · 统计学 2021-07-02 Yang Feng , Min Zhou , Xin Tong

With their increasing size, large language models (LLMs) are becoming increasingly good at language understanding tasks. But even with high performance on specific downstream task, LLMs fail at simple linguistic tests for negation or…

计算与语言 · 计算机科学 2023-12-01 Akshat Gupta

We derive a Bell-type inequality for observables with arbitrary spectra. For the case of continuous variable systems we propose a possible experimental violation of this inequality, by using squeezed light and homodyne detection together…

量子物理 · 物理学 2009-02-24 E. Shchukin W. Vogel

Hypothesis tests under order restrictions arise in a wide range of scientific applications. By exploiting inequality constraints, such tests can achieve substantial gains in power and interpretability. However, these gains come at a cost:…

统计方法学 · 统计学 2026-02-18 Ori Davidov

The ability to accurately estimate the sample size required by a stepped-wedge (SW) cluster randomized trial (CRT) routinely depends upon the specification of several nuisance parameters. If these parameters are mis-specified, the trial…

统计方法学 · 统计学 2017-10-10 Michael Grayling , Adrian Mander , James Wason

Large Language Models (LLMs) are increasingly used to evaluate information retrieval (IR) systems, generating relevance judgments traditionally made by human assessors. Recent empirical studies suggest that LLM-based evaluations often align…