English
Related papers

Related papers: Sets of Half-Average Nulls Generate Risk-Limiting …

200 papers

The pursuit of robot generalists, agents capable of performing diverse tasks across diverse environments, demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrained: it is…

Current question-answering benchmarks predominantly focus on accuracy in realizable prediction tasks. Conditioned on a question and answer-key, does the most likely token match the ground truth? Such benchmarks necessarily fail to evaluate…

Machine Learning · Computer Science 2024-12-02 André F. Cruz , Moritz Hardt , Celestine Mendler-Dünner

Scoring rules promote rational and honest decision-making, which is important for model evaluation and becoming increasingly important for automated procedures such as `AutoML'. In this paper we survey common squared and logarithmic scoring…

Statistics Theory · Mathematics 2025-06-03 Raphael Sonabend , John Zobolas , Riccardo Be Bin , Philipp Kopper , Lukas Burk , Andreas Bender

This paper proposes novel tests for the absence of jumps in a univariate semimartingale and for the absence of common jumps in a bivariate semimartingale. Our methods rely on ratio statistics of power variations based on irregular…

Statistics Theory · Mathematics 2017-12-21 Ole Martin , Mathias Vetter

Label-free reinforcement learning enables large language models to improve reasoning capabilities without ground-truth supervision, typically by treating majority-voted answers as pseudo-labels. However, we identify a critical failure mode:…

Computation and Language · Computer Science 2026-03-24 Teng Pan , Yuchen Yan , Zixuan Wang , Ruiqing Zhang , Guiyang Hou , Wenqi Zhang , Weiming Lu , Jun Xiao , Yongliang Shen

As Large Language Models (LLMs) are integrated into various sectors, ensuring their reliability and safety is crucial. This necessitates rigorous probing and auditing to maintain their effectiveness and trustworthiness in practical…

Artificial Intelligence · Computer Science 2024-06-19 Maryam Amirizaniani , Elias Martin , Tanya Roosta , Aman Chadha , Chirag Shah

Semi-supervised learning (SSL) can reduce the need for large labelled datasets by incorporating unlabelled data into the training. This is particularly interesting for semantic segmentation, where labelling data is very costly and…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Sebastian Scherer , Robin Schön , Rainer Lienhart

Online freelance marketplaces, a rapidly growing part of the global labor market, are creating a fair environment where professional skills are the main factor for hiring. While these platforms can reduce bias from traditional hiring, the…

Human-Computer Interaction · Computer Science 2025-10-17 Wugeng Zheng , Guohou Shan

A canonical problem in social choice is how to aggregate ranked votes: given $n$ voters' rankings over $m$ candidates, what voting rule $f$ should we use to aggregate these votes into a single winner? One standard method for comparing…

Computer Science and Game Theory · Computer Science 2023-08-08 Bailey Flanigan , Daniel Halpern , Alexandros Psomas

NLP benchmarks have largely focused on short texts, such as sentences and paragraphs, even though long texts comprise a considerable amount of natural language in the wild. We introduce SCROLLS, a suite of tasks that require reasoning over…

Computation and Language · Computer Science 2022-10-13 Uri Shaham , Elad Segal , Maor Ivgi , Avia Efrat , Ori Yoran , Adi Haviv , Ankit Gupta , Wenhan Xiong , Mor Geva , Jonathan Berant , Omer Levy

We propose a methodology to construct tests for the null hypothesis that the pricing errors of a panel of asset returns are jointly equal to zero in a linear factor asset pricing model -- that is, the null of "zero alpha". We consider, as a…

Econometrics · Economics 2026-05-12 Daniele Massacci , Lucio Sarno , Lorenzo Trapani , Pierluigi Vallarino

Large language models (LLMs) can often accurately describe probability distributions using natural language, yet they still struggle to generate faithful samples from them. This mismatch limits their use in tasks requiring reliable…

Machine Learning · Computer Science 2026-04-24 Tim Z. Xiao , Johannes Zenn , Zhen Liu , Weiyang Liu , Robert Bamler , Bernhard Schölkopf

Static Application Security Testing (SAST) tools are integral to modern software development, yet their adoption is undermined by excessive false positives that weaken developer trust and demand costly manual triage. We present ZeroFalse, a…

Semi-supervised learning (SSL) enables prediction with limited labels, but high-stakes tabular applications (medical, credit, recidivism) require statistical fairness guarantees. We identify a structural conflict in tabular fair SSL through…

Machine Learning · Computer Science 2026-05-19 Hangchun Liang , Changchun Li

Large language models (LLMs) frequently generate multiple candidate responses for a given prompt, yet selecting the most reliable one remains challenging, especially when correctness diverges from surface-level majority agreement. Existing…

Computation and Language · Computer Science 2026-04-15 Manh Nguyen , Sunil Gupta , Hung Le

We consider clinical trials with multiple, overlapping patient populations, that test multiple treatment policies specifically tailored to these populations. Such designs may lead to multiplicity issues, as false statements will affect…

Methodology · Statistics 2025-11-13 Remi Luschei , Werner Brannath

Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but their evaluations often collapse behavior into final task success. AgentAtlas reframes agent evaluation as a…

Artificial Intelligence · Computer Science 2026-05-27 Parsa Mazaheri , Kasra Mazaheri

In this paper, we present a novel algorithm to solve the Boolean Satisfiability (SAT) problem, using noise-based logic (NBL). Contrary to what the name may suggest, NBL is not a random/fuzzy logic system. In fact, it is a completely…

Computational Complexity · Computer Science 2011-10-05 Pey-Chang Kent Lin , Ayan Mandal , Sunil P Khatri

This paper extends the link between stochastic approximation (SA) theory and randomized urn models developed in Laruelle, Pag{\`e}s (2013), and their applications to clinical trials introduced in Bai, HU (1999,2005) and Bai, Hu, Shen…

Probability · Mathematics 2018-05-16 Sophie Laruelle , Gilles Pagès

We introduce the anytime-valid (AV) logrank test, a version of the logrank test that provides type-I error guarantees under optional stopping and optional continuation. The test is sequential without the need to specify a maximum sample…

Methodology · Statistics 2023-05-02 J. ter Schure , M. F. Perez-Ortiz , A. Ly , P. Grunwald
‹ Prev 1 8 9 10 Next ›