English
Related papers

Related papers: Degrees of Equivalence in a Key Comparison

200 papers

Misalignment between claims and their cited evidence is a common failure mode in reports generated by large language models, limiting their reliability in scientific and other high-stakes settings. We present DeepSciVerify, a two-stage…

Artificial Intelligence · Computer Science 2026-05-28 Shaghayegh Sadeghi , Khashayar Khajavi , Rise Adhikari , Alexander Tessier

There is a well-known problem in Null Hypothesis Significance Testing: many statistically significant results fail to replicate in subsequent experiments. We show that this problem arises because standard `point-form null' significance…

Methodology · Statistics 2025-02-06 Fintan Costello , Paul Watts

Large Language Models (LLMs) are increasingly utilized for domain-specific tasks, yet evaluating their outputs remains challenging. A common strategy is to apply evaluation criteria to assess alignment with domain-specific standards, yet…

Human-Computer Interaction · Computer Science 2026-02-17 Annalisa Szymanski , Simret Araya Gebreegziabher , Oghenemaro Anuyah , Ronald A. Metoyer , Toby Jia-Jun Li

This paper provides a general framework for testing instrument validity in heterogeneous causal effect models. The generalization includes the cases where the treatment can be multivalued ordered or unordered. Based on a series of testable…

Econometrics · Economics 2023-10-11 Zhenting Sun

Despite the growing interest in jailbreak methods as an effective red-teaming tool for building safe and responsible large language models (LLMs), flawed evaluation system designs have led to significant discrepancies in their effectiveness…

Computation and Language · Computer Science 2025-07-10 Ruixuan Huang , Xunguang Wang , Zongjie Li , Daoyuan Wu , Shuai Wang

In recent years, semantic similarity measure has a great interest in Semantic Web and Natural Language Processing (NLP). Several similarity measures have been developed, being given the existence of a structured knowledge representation…

Computation and Language · Computer Science 2013-10-31 Thabet Slimani

Vector similarity measures play a fundamental role in various fields, including machine learning, natural language processing, information retrieval, and data mining. These measures quantify the closeness between two vectors in a…

General Mathematics · Mathematics 2025-05-01 Abeeb A. Awotunde

Clinical language processing has received a lot of attention in recent years, resulting in new models or methods for disease phenotyping, mortality prediction, and other tasks. Unfortunately, many of these approaches are tested under…

Computation and Language · Computer Science 2022-09-30 Travis R. Goodwin , Dina Demner-Fushman

In this thesis, I refine our understanding as to what conclusions we can reach from coreference-based evaluations by expanding existing evaluation practices and considering the extent to which evaluation results are either converging or…

Computation and Language · Computer Science 2026-02-19 Ian Porada

Bell tests---the experimental demonstration of a Bell inequality violation---are central to understanding the foundations of quantum mechanics, underpin quantum technologies, and are a powerful diagnostic tool for technological developments…

The effective assessment of the instruction-following ability of large language models (LLMs) is of paramount importance. A model that cannot adhere to human instructions might be not able to provide reliable and helpful responses. In…

Computation and Language · Computer Science 2023-11-17 Yimin Jing , Renren Jin , Jiahao Hu , Huishi Qiu , Xiaohua Wang , Peng Wang , Deyi Xiong

The use of the fundamental constants for the definition of the most important measurement units of the International System, was considered a good solution to found it on more solid bases. From further analysis, this solution was found…

General Physics · Physics 2018-11-09 Franco Pavese

Large Vision-Language Models (LVLMs) suffer from hallucination issues, wherein the models generate plausible-sounding but factually incorrect outputs, undermining their reliability. A comprehensive quantitative evaluation is necessary to…

Computation and Language · Computer Science 2024-10-07 Haoyi Qiu , Wenbo Hu , Zi-Yi Dou , Nanyun Peng

Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output ("model choice") as a single data point underutilizes the information…

Artificial Intelligence · Computer Science 2025-12-30 Hongshen Sun , Juanjuan Zhang

The increasing use of text as data in social science research necessitates the development of valid, consistent, reproducible, and efficient methods for generating text-based concept measures. This paper presents a novel method that…

Computation and Language · Computer Science 2024-09-20 Yi Yang , Hanyu Duan , Jiaxin Liu , Kar Yan Tam

The ability to compare the semantic similarity between text corpora is important in a variety of natural language processing applications. However, standard methods for evaluating these metrics have yet to be established. We propose a set…

Computation and Language · Computer Science 2022-11-30 George Kour , Samuel Ackerman , Orna Raz , Eitan Farchi , Boaz Carmeli , Ateret Anaby-Tavor

Proximities are at the heart of almost all machine learning methods. If the input data are given as numerical vectors of equal lengths, euclidean distance, or a Hilbertian inner product is frequently used in modeling algorithms. In a more…

Machine Learning · Computer Science 2020-09-01 Maximilian Münch , Michiel Straat , Michael Biehl , Frank-Michael Schleif

Research accomplishment is usually measured by considering all citations with equal importance, thus ignoring the wide variety of purposes an article is being cited for. Here, we posit that measuring the intensity of a reference is crucial…

Computation and Language · Computer Science 2016-09-02 Tanmoy Chakraborty , Ramasuri Narayanam

A new method, with an application program in Matlab code, is proposed for testing item performance models on empirical databases. This method uses data intraclass correlation statistics as expected correlations to which one compares simple…

Several performance measures can be used for evaluating classification results: accuracy, F-measure, and many others. Can we say that some of them are better than others, or, ideally, choose one measure that is best in all situations? To…

Machine Learning · Computer Science 2022-01-25 Martijn Gösgens , Anton Zhiyanov , Alexey Tikhonov , Liudmila Prokhorenkova