中文
相关论文

相关论文: Inference at Scale Significance Testing for Large …

200 篇论文

Test-time scaling (TTS) has emerged as a new frontier for scaling the performance of Large Language Models. In test-time scaling, by using more computational resources during inference, LLMs can improve their reasoning process and task…

计算与语言 · 计算机科学 2025-09-10 V Venktesh , Mandeep Rathee , Avishek Anand

Relevance evaluation of a query and a passage is essential in Information Retrieval (IR). Recently, numerous studies have been conducted on tasks related to relevance judgment using Large Language Models (LLMs) such as GPT-4, demonstrating…

信息检索 · 计算机科学 2024-05-14 Jaekeol Choi

Evaluation of reasoning language models gained importance after it was observed that they can combine their existing capabilities into novel traces of intermediate steps before task completion and that the traces can sometimes help them to…

机器学习 · 计算机科学 2025-08-15 Petr Spelda , Vit Stritecky

An exciting recent development is the uptake of deep neural networks in many scientific fields, where the main objective is outcome prediction with the black-box nature. Significance testing is promising to address the black-box issue and…

机器学习 · 统计学 2022-06-22 Ben Dai , Xiaotong Shen , Wei Pan

Incomplete relevance judgments limit the re-usability of test collections. When new systems are compared against previous systems used to build the pool of judged documents, they often do so at a disadvantage due to the ``holes'' in test…

信息检索 · 计算机科学 2024-05-10 Zahra Abbasiantaeb , Chuan Meng , Leif Azzopardi , Mohammad Aliannejadi

This paper deals with the issue of testing hypothesis in symmetric and log-symmetric linear regression models in small and moderate-sized samples. We focus on four tests, namely the Wald, likelihood ratio, score, and gradient tests. These…

统计方法学 · 统计学 2016-02-03 Francisco M. C. Medeiros , Silvia L. P. Ferrari

Tests based on heteroskedasticity robust standard errors are an important technique in econometric practice. Choosing the right critical value, however, is not simple at all: conventional critical values based on asymptotics often lead to…

统计理论 · 数学 2025-05-07 Benedikt M. Pötscher , David Preinerstorfer

Motivation: P values derived from the null hypothesis significance testing framework are strongly affected by sample size, and are known to be irreproducible in underpowered studies, yet no suitable replacement has been proposed. Results:…

In observational studies, researchers must select a method to control for confounding. Options include propensity score methods and regression. It remains unclear how dataset characteristics (size, overlap in propensity scores, exposure…

统计方法学 · 统计学 2022-10-21 J. Wilkinson , M. A. Mamas , E. Kontopantelis

In this paper, we develop invariance-based procedures for testing and inference in high-dimensional regression models. These procedures, also known as randomization tests, provide several important advantages. First, for the global null…

统计方法学 · 统计学 2023-12-27 Wenxuan Guo , Panos Toulis

Many networks in natural and human-made systems exhibit scale-free properties and are small worlds. Now we show that people's understanding of complex systems in their cognitive maps also follow a scale-free topology (P_k = k^-lambda,…

神经元与认知 · 定量生物学 2007-05-23 Uygar Ozesmi , Can Ozan Tan

Human relevance assessment is time-consuming and cognitively intensive, limiting the scalability of Information Retrieval evaluation. This has led to growing interest in using large language models (LLMs) as proxies for human judges.…

信息检索 · 计算机科学 2026-04-28 Chuting Yu , Hang Li , Guido Zuccon , Joel Mackenzie , Teerapong Leelanupab

Instruction subsets are heuristics that can reduce the size of the inductive programming search space by tens of orders of magnitude. Comprising many overlapping subsets of different sizes, they serve as predictions of the instructions…

人工智能 · 计算机科学 2024-07-02 Edward McDaid , Sarah McDaid

Testing of deep learning models is challenging due to the excessive number and complexity of computations involved. As a result, test data selection is performed manually and in an ad hoc way. This raises the question of how we can…

机器学习 · 计算机科学 2019-05-01 Wei Ma , Mike Papadakis , Anestis Tsakmalis , Maxime Cordy , Yves Le Traon

Large Reasoning Models (LRMs) achieve strong performance on mathematical reasoning tasks but remain unreliable on challenging instances. Existing test-time scaling methods, such as repeated sampling, self-correction, and tree search,…

人工智能 · 计算机科学 2026-04-30 Zhimin Lin , Yixin Ji , Jinpeng Li , Yu Luo , Dong Li , Junhua Fang , Juntao Li , Min Zhang

When a scientist performs an experiment they normally acquire a set of measurements and are expected to demonstrate that their results are "statistically significant" thus confirming whatever hypothesis they are testing. The main method for…

其他统计学 · 统计学 2011-09-30 Jacob Levman

Rank-based inference methods are applied in various disciplines, typically when procedures relying on standard normal theory are not justifiable, for example when data are not symmetrically distributed, contain outliers, or responses are…

统计理论 · 数学 2018-02-16 Edgar Brunner , Frank Konietschke , Arne C. Bathke , Markus Pauly

To assess whether there is some signal in a big database, aggregate tests for the global null hypothesis of no effect are routinely applied in practice before more specialized analysis is carried out. Although a plethora of aggregate tests…

统计理论 · 数学 2024-05-08 Anders Bredahl Kock , David Preinerstorfer

Small study effects occur when smaller studies show different, often larger, treatment effects than large ones, which may threaten the validity of systematic reviews and meta-analyses. The most well-known reasons for small study effects…

统计方法学 · 统计学 2018-05-28 Chuan Hong , Georgia Salanti , Sally Morton , Richard Riley , Haitao Chu , Stephen E. Kimmel , Yong Chen

Large language models (LLMs) obtain state of the art zero shot relevance ranking performance on a variety of information retrieval tasks. The two most common prompts to elicit LLM relevance judgments are pointwise scoring (a.k.a. relevance…

机器学习 · 计算机科学 2025-05-27 Charles Godfrey , Ping Nie , Natalia Ostapuk , David Ken , Shang Gao , Souheil Inati