中文
相关论文

相关论文: Inference at Scale Significance Testing for Large …

200 篇论文

Large language models (LLMs) and multimodal LLMs (MLL-Ms) excel at chain-of-thought reasoning but face distribution shift at test-time and a lack of verifiable supervision. Recent test-time reinforcement learning (TTRL) methods derive…

计算与语言 · 计算机科学 2026-03-09 Jianghao Wu , Yasmeen George , Jin Ye , Yicheng Wu , Daniel F. Schmidt , Jianfei Cai

Transformer-based pre-trained models have recently achieved great results in solving many software engineering tasks including automatic code completion which is a staple in a developer's toolkit. While many have striven to improve the…

计算与语言 · 计算机科学 2023-04-25 Tim van Dam , Maliheh Izadi , Arie van Deursen

In evaluation campaigns, participants often explore variations of popular, state-of-the-art baselines as a low-risk strategy to achieve competitive results. While effective, this can lead to local "hill climbing" rather than more radical…

信息检索 · 计算机科学 2023-01-31 Mehmet Deniz Türkmen , Matthew Lease , Mucahid Kutlu

Big Data often presents as massive non-probability samples. Not only is the selection mechanism often unknown, but larger data volume amplifies the relative contribution of selection bias to total error. Existing bias adjustment approaches…

统计方法学 · 统计学 2022-03-29 Ali Rafei , Carol A. C. Flannagan , Brady T. West , Michael R. Elliott

To overcome the limitations of automated metrics (e.g. BLEU, METEOR) for evaluating dialogue systems, researchers typically use human judgments to provide convergent evidence. While it has been demonstrated that human judgments can suffer…

计算与语言 · 计算机科学 2019-09-24 Sashank Santhanam , Samira Shaikh

While the scaling laws of large language models (LLMs) training have been extensively studied, optimal inference configurations of LLMs remain underexplored. We study inference scaling laws (aka test-time scaling laws) and compute-optimal…

人工智能 · 计算机科学 2025-03-04 Yangzhen Wu , Zhiqing Sun , Shanda Li , Sean Welleck , Yiming Yang

Despite its importance to experimental design, statistical power (the probability that, given a real effect, an experiment will reject the null hypothesis) has largely been ignored by the NLP community. Underpowered experiments make it more…

计算与语言 · 计算机科学 2020-10-15 Dallas Card , Peter Henderson , Urvashi Khandelwal , Robin Jia , Kyle Mahowald , Dan Jurafsky

Traditional evaluation of information retrieval (IR) systems relies on human-annotated relevance labels, which can be both biased and costly at scale. In this context, large language models (LLMs) offer an alternative by allowing us to…

信息检索 · 计算机科学 2024-10-21 Naghmeh Farzi , Laura Dietz

Sphericity test plays a key role in many statistical problems. We propose Spearman's rho-type rank test and Kendall's tau-type rank test for sphericity in the high dimensional settings. We show that these two tests are equivalent. Thanks to…

统计方法学 · 统计学 2015-02-17 Long Feng

In-context learning (ICL), the ability of large language models to perform novel tasks by conditioning on a prompt with a few task examples, requires these examples to be informative about the test instance. The standard approach of…

计算与语言 · 计算机科学 2023-11-08 Shivanshu Gupta , Matt Gardner , Sameer Singh

Large language models (LLMs) have shown an impressive ability to perform a wide range of tasks using in-context learning (ICL), where a few examples are used to describe a task to the model. However, the performance of ICL varies…

计算与语言 · 计算机科学 2024-06-25 Keqin Peng , Liang Ding , Yancheng Yuan , Xuebo Liu , Min Zhang , Yuanxin Ouyang , Dacheng Tao

Graph-constrained estimation methods encourage similarities among neighboring covariates presented as nodes on a graph, which can result in more accurate estimations, especially in high dimensional settings. Variable selection approaches…

统计方法学 · 统计学 2018-05-29 Sen Zhao , Ali Shojaie

In this paper, we propose a power comparison between high dimensional t-test, sign and signed rank test for the one sample mean test. We show that the high dimensional signed rank test is superior to a high dimensional t test, but inferior…

统计方法学 · 统计学 2018-12-31 Long Feng

Experimental comparisons of performance represent an important aspect of research on optimization algorithms. In this work we present a methodology for defining the required sample sizes for designing experiments with desired statistical…

神经与进化计算 · 计算机科学 2018-10-16 Felipe Campelo , Fernanda Takahashi

Testing fit of data to model is fundamentally important to any science, but publications in the field of phylogenetics rarely do this. Such analyses discard fundamental aspects of science as prescribed by Karl Popper. Indeed, not without…

种群与进化 · 定量生物学 2008-12-31 Peter J. Waddell , Rissa Ota , David Penny

A fundamental challenge in comparing two survival distributions with right censored data is the selection of an appropriate nonparametric test, as the power of standard tests like the Log rank and Wilcoxon is highly dependent on the often…

统计方法学 · 统计学 2025-10-09 Abid Hussain , Touqeer Ahmad

Most statistical tests for treatment effects used in randomized clinical trials with survival outcomes are based on the proportional hazards assumption, which often fails in practice. Data from early exploratory studies may provide evidence…

统计理论 · 数学 2020-05-28 Andrea Arfé , Brian Alexander , Lorenzo Trippa

How should researchers analyze randomized experiments in which the main outcome is latent and measured in multiple ways but each measure contains some degree of error? We first identify a critical study-specific noncomparability problem in…

计量经济学 · 经济学 2026-01-13 Jiawei Fu , Donald P. Green

Most of the statistical tests currently used to detect differentially expressed genes are based on asymptotic results, and perform poorly for low expression tags. Another problem is the common use of a single canonical cutoff for the…

基因组学 · 定量生物学 2008-08-04 Leonardo Varuzza , Arthur Gruber , Carlos A. de B. Pereira

Large Language Models (LLMs) are increasingly used to automate relevance judgments for information retrieval (IR) tasks, often demonstrating agreement with human labels that approaches inter-human agreement. To assess the robustness and…

信息检索 · 计算机科学 2025-04-18 Negar Arabzadeh , Charles L. A . Clarke