中文
相关论文

相关论文: Faithful Model Evaluation for Model-Based Metrics

200 篇论文

Positive predictive value and negative predictive value are two widely used parameters to assess the clinical usefulness of a medical diagnostic test. When there are two diagnostic tests, it is recommendable to make a comparative assessment…

统计方法学 · 统计学 2024-05-29 Antonio Martín Andrés , Pedro Femia Marzo

Current evaluation metrics for language modeling and generation rely heavily on the accuracy of predicted (or generated) words as compared to a reference ground truth. While important, token-level accuracy only captures one aspect of a…

计算与语言 · 计算机科学 2020-10-15 Shiran Dudy , Steven Bedrick

Importance sampling is a common technique for Monte Carlo approximation, including Monte Carlo approximation of p-values. Here it is shown that a simple correction of the usual importance sampling p-values creates valid p-values, meaning…

统计计算 · 统计学 2011-04-12 Matthew T. Harrison

Recently, it was shown that most popular IR measures are not interval-scaled, implying that decades of experimental IR research used potentially improper methods, which may have produced questionable results. However, it was unclear if and…

信息检索 · 计算机科学 2021-01-08 Marco Ferrante , Nicola Ferro , Norbert Fuhr

Large language models (LLMs) achieve strong performance and have revolutionized NLP, but their lack of explainability keeps them treated as black boxes, limiting their use in domains that demand transparency and trust. A promising direction…

计算与语言 · 计算机科学 2026-04-17 Bar Alon , Itamar Zimerman , Lior Wolf

When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' growing capabilities, however there has been no similar focus…

机器学习 · 计算机科学 2025-02-06 Joshua Vendrow , Edward Vendrow , Sara Beery , Aleksander Madry

Large language models (LLMs) are increasingly used to support the analysis of complex financial disclosures, yet their reliability, behavioral consistency, and transparency remain insufficiently understood in high-stakes settings. This…

计算与语言 · 计算机科学 2026-01-21 Md Talha Mohsin

We introduce a method to measure uncertainty in large language models. For tasks like question answering, it is essential to know when we can trust the natural language outputs of foundation models. We show that measuring uncertainty in…

计算与语言 · 计算机科学 2023-04-18 Lorenz Kuhn , Yarin Gal , Sebastian Farquhar

The logical and practical difficulties associated with research interpretation using P values and null hypothesis significance testing have been extensively documented. This paper describes an alternative, likelihood-based approach to…

统计方法学 · 统计学 2021-09-21 Nicholas Adams , Gerard O'Reilly

The proliferation of Large Language Models (LLMs) is challenged by hallucinations, critical failure modes where models generate non-factual, nonsensical or unfaithful text. This paper introduces Semantic Divergence Metrics (SDM), a novel…

计算与语言 · 计算机科学 2025-08-15 Igor Halperin

With the widespread application of Large Language Models (LLMs) to various domains, concerns regarding the trustworthiness of LLMs in safety-critical scenarios have been raised, due to their unpredictable tendency to hallucinate and…

计算与语言 · 计算机科学 2024-11-04 Xin Qiu , Risto Miikkulainen

Uncertainty quantification is a set of techniques that measure confidence in language models. They can be used, for example, to detect hallucinations or alert users to review uncertain predictions. To be useful, these confidence scores must…

计算与语言 · 计算机科学 2026-04-13 Lorenzo Jaime Yu Flores , Cesare Spinoso di-Piano , Jackie Chi Kit Cheung

Large language models (LLMs) are increasingly used as decision-support tools in data-constrained scientific workflows, where correctness and validity are critical. However, evaluation practices often emphasize stability or reproducibility…

机器学习 · 计算机科学 2026-03-18 Nazia Riasat

Recent advancements in large language models have led to significant improvements across various tasks, including mathematical reasoning, which is used to assess models' intelligence in logical reasoning and problem-solving. Models are…

人工智能 · 计算机科学 2026-04-27 Erez Yosef , Oron Anschel , Shunit Haviv Hakimi , Asaf Gendler , Adam Botach , Nimrod Berman , Igor Kviatkovsky

We examine the role of trustworthiness and trust in statistical inference, arguing that it is the extent of trustworthiness in inferential statistical tools which enables trust in the conclusions. Certain tools, such as the p-value and…

统计方法学 · 统计学 2021-05-11 David J. Hand

Large Language Models (LLMs) are capable of generating persuasive Natural Language Explanations (NLEs) to justify their answers. However, the faithfulness of these explanations should not be readily trusted at face value. Recent studies…

计算与语言 · 计算机科学 2024-11-04 Wei Jie Yeo , Ranjan Satapathy , Erik Cambria

When a scientist performs an experiment they normally acquire a set of measurements and are expected to demonstrate that their results are "statistically significant" thus confirming whatever hypothesis they are testing. The main method for…

其他统计学 · 统计学 2011-09-30 Jacob Levman

The staggering pace with which the capabilities of large language models (LLMs) are increasing, as measured by a range of commonly used natural language understanding (NLU) benchmarks, raises many questions regarding what "understanding"…

计算与语言 · 计算机科学 2024-04-19 Xenia Ohmer , Elia Bruni , Dieuwke Hupkes

We describe a modified sequential probability ratio test that can be used to reduce the average sample size required to perform statistical hypothesis tests at specified levels of significance and power. Examples are provided for $z$ tests,…

统计方法学 · 统计学 2020-12-04 Sandipan Pramanik , Valen E. Johnson , Anirban Bhattacharya

As Large Language Models and Natural Language Processing (NLP) technology rapidly develop and spread into daily life, it becomes crucial to anticipate how their use could harm people. One problem that has received a lot of attention in…

计算与语言 · 计算机科学 2024-01-17 Oskar van der Wal , Dominik Bachmann , Alina Leidinger , Leendert van Maanen , Willem Zuidema , Katrin Schulz