中文
相关论文

相关论文: Sets of Half-Average Nulls Generate Risk-Limiting …

200 篇论文

The pursuit of robot generalists, agents capable of performing diverse tasks across diverse environments, demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrained: it is…

Current question-answering benchmarks predominantly focus on accuracy in realizable prediction tasks. Conditioned on a question and answer-key, does the most likely token match the ground truth? Such benchmarks necessarily fail to evaluate…

机器学习 · 计算机科学 2024-12-02 André F. Cruz , Moritz Hardt , Celestine Mendler-Dünner

Scoring rules promote rational and honest decision-making, which is important for model evaluation and becoming increasingly important for automated procedures such as `AutoML'. In this paper we survey common squared and logarithmic scoring…

This paper proposes novel tests for the absence of jumps in a univariate semimartingale and for the absence of common jumps in a bivariate semimartingale. Our methods rely on ratio statistics of power variations based on irregular…

统计理论 · 数学 2017-12-21 Ole Martin , Mathias Vetter

Label-free reinforcement learning enables large language models to improve reasoning capabilities without ground-truth supervision, typically by treating majority-voted answers as pseudo-labels. However, we identify a critical failure mode:…

计算与语言 · 计算机科学 2026-03-24 Teng Pan , Yuchen Yan , Zixuan Wang , Ruiqing Zhang , Guiyang Hou , Wenqi Zhang , Weiming Lu , Jun Xiao , Yongliang Shen

As Large Language Models (LLMs) are integrated into various sectors, ensuring their reliability and safety is crucial. This necessitates rigorous probing and auditing to maintain their effectiveness and trustworthiness in practical…

人工智能 · 计算机科学 2024-06-19 Maryam Amirizaniani , Elias Martin , Tanya Roosta , Aman Chadha , Chirag Shah

Semi-supervised learning (SSL) can reduce the need for large labelled datasets by incorporating unlabelled data into the training. This is particularly interesting for semantic segmentation, where labelling data is very costly and…

计算机视觉与模式识别 · 计算机科学 2022-10-20 Sebastian Scherer , Robin Schön , Rainer Lienhart

Online freelance marketplaces, a rapidly growing part of the global labor market, are creating a fair environment where professional skills are the main factor for hiring. While these platforms can reduce bias from traditional hiring, the…

人机交互 · 计算机科学 2025-10-17 Wugeng Zheng , Guohou Shan

A canonical problem in social choice is how to aggregate ranked votes: given $n$ voters' rankings over $m$ candidates, what voting rule $f$ should we use to aggregate these votes into a single winner? One standard method for comparing…

计算机科学与博弈论 · 计算机科学 2023-08-08 Bailey Flanigan , Daniel Halpern , Alexandros Psomas

NLP benchmarks have largely focused on short texts, such as sentences and paragraphs, even though long texts comprise a considerable amount of natural language in the wild. We introduce SCROLLS, a suite of tasks that require reasoning over…

计算与语言 · 计算机科学 2022-10-13 Uri Shaham , Elad Segal , Maor Ivgi , Avia Efrat , Ori Yoran , Adi Haviv , Ankit Gupta , Wenhan Xiong , Mor Geva , Jonathan Berant , Omer Levy

We propose a methodology to construct tests for the null hypothesis that the pricing errors of a panel of asset returns are jointly equal to zero in a linear factor asset pricing model -- that is, the null of "zero alpha". We consider, as a…

计量经济学 · 经济学 2026-05-12 Daniele Massacci , Lucio Sarno , Lorenzo Trapani , Pierluigi Vallarino

Large language models (LLMs) can often accurately describe probability distributions using natural language, yet they still struggle to generate faithful samples from them. This mismatch limits their use in tasks requiring reliable…

机器学习 · 计算机科学 2026-04-24 Tim Z. Xiao , Johannes Zenn , Zhen Liu , Weiyang Liu , Robert Bamler , Bernhard Schölkopf

Static Application Security Testing (SAST) tools are integral to modern software development, yet their adoption is undermined by excessive false positives that weaken developer trust and demand costly manual triage. We present ZeroFalse, a…

Semi-supervised learning (SSL) enables prediction with limited labels, but high-stakes tabular applications (medical, credit, recidivism) require statistical fairness guarantees. We identify a structural conflict in tabular fair SSL through…

机器学习 · 计算机科学 2026-05-19 Hangchun Liang , Changchun Li

Large language models (LLMs) frequently generate multiple candidate responses for a given prompt, yet selecting the most reliable one remains challenging, especially when correctness diverges from surface-level majority agreement. Existing…

计算与语言 · 计算机科学 2026-04-15 Manh Nguyen , Sunil Gupta , Hung Le

We consider clinical trials with multiple, overlapping patient populations, that test multiple treatment policies specifically tailored to these populations. Such designs may lead to multiplicity issues, as false statements will affect…

统计方法学 · 统计学 2025-11-13 Remi Luschei , Werner Brannath

Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but their evaluations often collapse behavior into final task success. AgentAtlas reframes agent evaluation as a…

人工智能 · 计算机科学 2026-05-27 Parsa Mazaheri , Kasra Mazaheri

In this paper, we present a novel algorithm to solve the Boolean Satisfiability (SAT) problem, using noise-based logic (NBL). Contrary to what the name may suggest, NBL is not a random/fuzzy logic system. In fact, it is a completely…

计算复杂性 · 计算机科学 2011-10-05 Pey-Chang Kent Lin , Ayan Mandal , Sunil P Khatri

This paper extends the link between stochastic approximation (SA) theory and randomized urn models developed in Laruelle, Pag{\`e}s (2013), and their applications to clinical trials introduced in Bai, HU (1999,2005) and Bai, Hu, Shen…

概率论 · 数学 2018-05-16 Sophie Laruelle , Gilles Pagès

We introduce the anytime-valid (AV) logrank test, a version of the logrank test that provides type-I error guarantees under optional stopping and optional continuation. The test is sequential without the need to specify a maximum sample…

统计方法学 · 统计学 2023-05-02 J. ter Schure , M. F. Perez-Ortiz , A. Ly , P. Grunwald
‹ 上一页 1 8 9 10 下一页 ›