English
Related papers

Related papers: Interlaboratory consensus building challenge

200 papers

Performance in cross-lingual NLP tasks is impacted by the (dis)similarity of languages at hand: e.g., previous work has suggested there is a connection between the expected success of bilingual lexicon induction (BLI) and the assumption of…

Computation and Language · Computer Science 2020-10-13 Haim Dubossarsky , Ivan Vulić , Roi Reichart , Anna Korhonen

Contemporary approaches to assisted scientific discovery use language models to automatically generate large numbers of potential hypothesis to test, while also automatically generating code-based experiments to test those hypotheses. While…

Artificial Intelligence · Computer Science 2025-09-23 Peter Jansen , Samiah Hassan , Ruoyao Wang

Link prediction is a paradigmatic problem in network science, which aims at estimating the existence likelihoods of nonobserved links, based on known topology. After a brief introduction of the standard problem and metrics of link…

Data Analysis, Statistics and Probability · Physics 2021-12-14 Tao Zhou

Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep epistemic divergence. Using two major reasoning benchmarks -…

Computation and Language · Computer Science 2026-02-13 Eddie Yang , Dashun Wang

Natural Language Inference (NLI) is the task of determining whether a premise entails, contradicts, or is neutral with respect to a given hypothesis. The task is often framed as emulating human inferential processes, in which commonsense…

Computation and Language · Computer Science 2026-01-27 Chathuri Jayaweera , Brianna Yanqui , Bonnie Dorr

Despite the superior capabilities of Multimodal Large Language Models (MLLMs) across diverse tasks, they still face significant trustworthiness challenges. Yet, current literature on the assessment of trustworthy MLLMs remains limited,…

Computation and Language · Computer Science 2024-12-09 Yichi Zhang , Yao Huang , Yitong Sun , Chang Liu , Zhe Zhao , Zhengwei Fang , Yifan Wang , Huanran Chen , Xiao Yang , Xingxing Wei , Hang Su , Yinpeng Dong , Jun Zhu

The rapid rise of large language models (LLMs) is reshaping the landscape of automatic assessment in education. While these systems demonstrate substantial advantages in adaptability to diverse question types and flexibility in output…

Large language models (LLMs) are increasingly used for annotation in computational social science, yet their methodological reliability under prompt variation remains unclear. This paper introduces Inter-Prompt Reliability (IPR), a…

Computers and Society · Computer Science 2026-04-21 Jingyuan Liu

In the context of industrially mass-manufactured products, quality management is based on physically inspecting a small sample from a large batch and reasoning about the batch's quality conformance. When complementing physical inspections…

Applications · Statistics 2024-02-22 Simon Cramer , Tobias Müller , Robert H. Schmitt

Intent classification and slot filling are two critical tasks for natural language understanding. Traditionally the two tasks have been deemed to proceed independently. However, more recently, joint models for intent classification and slot…

Computation and Language · Computer Science 2021-02-23 H. Weld , X. Huang , S. Long , J. Poon , S. C. Han

While the NLP community has produced numerous summarization benchmarks, none provide the rich annotations required to simultaneously address many important problems related to control and reliability. We introduce a Wikipedia-derived…

Computation and Language · Computer Science 2023-12-05 Kundan Krishna , Prakhar Gupta , Sanjana Ramprasad , Byron C. Wallace , Jeffrey P. Bigham , Zachary C. Lipton

Large Language Models (LLMs) have shown impressive capabilities in various applications, but they still face various inconsistency issues. Existing works primarily focus on the inconsistency issues within a single LLM, while we…

Computation and Language · Computer Science 2024-11-15 Kai Xiong , Xiao Ding , Yixin Cao , Ting Liu , Bing Qin

Large language models (LLMs) produce outputs with varying levels of uncertainty, and, just as often, varying levels of correctness; making their practical reliability far from guaranteed. To quantify this uncertainty, we systematically…

Computation and Language · Computer Science 2025-10-24 Christian Hobelsberger , Theresa Winner , Andreas Nawroth , Oliver Mitevski , Anna-Carolina Haensch

Background: Philosophers of science including Collins, Feyerabend, Kuhn and Latour have all emphasized the importance of consensus within scientific communities of practice. Consensus is important for maintaining legitimacy with outsiders,…

Software Engineering · Computer Science 2018-02-20 Pontus Johnson , Paul Ralph , Mathias Ekstedt , Iaakov Exman , Michael Goedicke

The internal validity of observational study is often subject to debate. In this study, we define the unobserved sample based on the counterfactuals and formalize its relationship with the null hypothesis statistical testing (NHST) for…

Methodology · Statistics 2020-05-27 Tenglong Li , Kenneth A. Frank

Uncertainty estimation is important for ensuring safety and robustness of AI systems. While most research in the area has focused on un-structured prediction tasks, limited work has investigated general uncertainty estimation approaches for…

Machine Learning · Statistics 2021-02-12 Andrey Malinin , Mark Gales

Pairwise comparisons are an important tool of modern (multiple criteria) decision making. Since human judgments are often inconsistent, many studies focused on the ways how to express and measure this inconsistency, and several…

Artificial Intelligence · Computer Science 2017-05-01 Jiri Mazurek

We propose a collaborative framework in which multiple large language models -- including GPT-4-0125-preview, Meta-LLaMA-3-70B-Instruct, Claude-3-Opus, and Gemini-1.5-Flash -- generate and answer complex, PhD-level statistical questions…

Computation and Language · Computer Science 2025-02-25 Alireza Amiri-Margavi , Iman Jebellat , Ehsan Jebellat , Seyed Pouyan Mousavi Davoudi

Multiple raters are often needed to be used interchangeably in practice for measurement or evaluation. Assessing agreement among these multiple raters via agreement indices are necessary before their participation. While the intuitively…

Methodology · Statistics 2020-06-09 Tongrong Wang , Huiman X. Barnhart

Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by various failure modes that impact benchmark bias, variance,…

‹ Prev 1 8 9 10 Next ›