中文
相关论文

相关论文: NovAScore: A New Automated Metric for Evaluating D…

200 篇论文

Factuality evaluation of large language model (LLM) outputs requires decomposing text into discrete "atomic" facts. However, existing definitions of atomicity are underspecified, with empirical results showing high disagreement among…

人机交互 · 计算机科学 2025-09-03 Manuel Schmidt , Daniel A. Keim , Frederik L. Dennig

Lee, Walsh, and Wang (2015) - based on Uzzi, Mukherjee, Stringer, and Jones (2013) - and Wang, Veugelers, and Stephan (2017) proposed scores based on cited references (cited journals) data which can be used to measure the novelty of papers…

数字图书馆 · 计算机科学 2019-10-09 Lutz Bornmann , Alexander Tekles , Helena H. Zhang , Fred Y. Ye

Collecting human judgements is currently the most reliable evaluation method for natural language generation systems. Automatic metrics have reported flaws when applied to measure quality aspects of generated text and have been shown to…

计算与语言 · 计算机科学 2022-04-29 Thórhildur Thorleiksdóttir , Cedric Renggli , Nora Hollenstein , Ce Zhang

Existing benchmarks for summarization quality evaluation often lack diverse input scenarios, focus on narrowly defined dimensions (e.g., faithfulness), and struggle with subjective and coarse-grained annotation schemes. To address these…

计算与语言 · 计算机科学 2024-10-02 Yuho Lee , Taewon Yun , Jason Cai , Hang Su , Hwanjun Song

Sentence level novelty detection aims at reducing redundant sentences from a sentence list. In the task, sentences appearing later in the list with no new meanings are eliminated. Aiming at a better accuracy for detecting redundancy, this…

信息检索 · 计算机科学 2007-05-23 Le Zhao , Min Zhang , Shaoping Ma

Evaluating text summarization has been a challenging task in natural language processing (NLP). Automatic metrics which heavily rely on reference summaries are not suitable in many situations, while human evaluation is time-consuming and…

计算与语言 · 计算机科学 2024-07-02 Huyen Nguyen , Haihua Chen , Lavanya Pobbathi , Junhua Ding

Many text generation applications require the generated text to be factually consistent with input information. Automatic evaluation of factual consistency is challenging. Previous work has developed various metrics that often depend on…

计算与语言 · 计算机科学 2023-05-29 Yuheng Zha , Yichi Yang , Ruichen Li , Zhiting Hu

Peer review serves as a backbone of academic research, but in most AI conferences, the review quality is degrading as the number of submissions explodes. To reliably detect low-quality reviews, we define misinformed review points as either…

Generation novelty is a key indicator of an LLM's ability to generalize, yet measuring it against full pretraining corpora is computationally challenging. Existing evaluations often rely on lexical overlap, failing to detect paraphrased…

机器学习 · 计算机科学 2026-01-14 Philipp Davydov , Ameya Prabhu , Matthias Bethge , Elisa Nguyen , Seong Joon Oh

Recent advancements in the field of natural language generation have facilitated the use of large language models to assess the quality of generated text. Although these models have shown promising results in tasks such as machine…

人工智能 · 计算机科学 2024-01-23 Terry Yue Zhuo

The most prevalent scope of interest for OCR applications used to be scanned documents, but it has now shifted towards the natural scene. Despite the change of times, the existing evaluation methods are still based on the old criteria…

计算机视觉与模式识别 · 计算机科学 2019-08-30 Hong-Seok Lee , Youngmin Yoon , Pil-Hoon Jang , Chankyu Choi

Text summarization models are often trained to produce summaries that meet human quality requirements. However, the existing evaluation metrics for summary text are only rough proxies for summary quality, suffering from low correlation with…

计算与语言 · 计算机科学 2022-07-12 Wuhang Lin , Shasha Li , Chen Zhang , Bin Ji , Jie Yu , Jun Ma , Zibo Yi

We suggest a new method for creating and using gold-standard datasets for word similarity evaluation. Our goal is to improve the reliability of the evaluation, and we do this by redesigning the annotation task to achieve higher inter-rater…

计算与语言 · 计算机科学 2017-02-28 Oded Avraham , Yoav Goldberg

This paper proposes a framework for assessing the novelty of design problems using the SAPPhIRE model of causality. The novelty of a problem is measured as its minimum distance from the problems in a reference problem database. The distance…

计算与语言 · 计算机科学 2024-10-25 Sanjay Singh , Amaresh Chakrabarti

Current metrics for evaluating factuality for abstractive document summarization have achieved high correlations with human judgment, but they do not account for the vision modality and thus are not adequate for vision-and-language…

计算与语言 · 计算机科学 2022-11-07 David Wan , Mohit Bansal

Novel scientific knowledge is constantly produced by the scientific community. Understanding the level of novelty characterized by scientific literature is key for modeling scientific dynamics and analyzing the growth mechanisms of…

数字图书馆 · 计算机科学 2018-01-30 Jiangen He , Chaomei Chen

Existing automatic evaluation on text-to-image synthesis can only provide an image-text matching score, without considering the object-level compositionality, which results in poor correlation with human judgments. In this work, we propose…

计算机视觉与模式识别 · 计算机科学 2023-05-19 Yujie Lu , Xianjun Yang , Xiujun Li , Xin Eric Wang , William Yang Wang

The wide acceptance of large language models (LLMs) has unlocked new applications and social risks. Popular countermeasures aim at detecting misinformation, usually involve domain specific models trained to recognize the relevance of any…

计算与语言 · 计算机科学 2024-06-03 Edouard Yvinec , Gabriel Kasser

Multi-modal generative document parsing systems challenge traditional evaluation: unlike deterministic OCR or layout models, they often produce semantically correct yet structurally divergent outputs. Conventional metrics-CER, WER, IoU, or…

计算与语言 · 计算机科学 2025-09-25 Renyu Li , Antonio Jimeno Yepes , Yao You , Kamil Pluciński , Maximilian Operlejn , Crag Wolfe

Anthropomorphism, or the attribution of human-like characteristics to non-human entities, has shaped conversations about the impacts and possibilities of technology. We present AnthroScore, an automatic metric of implicit anthropomorphism…

计算与语言 · 计算机科学 2024-02-06 Myra Cheng , Kristina Gligoric , Tiziano Piccardi , Dan Jurafsky