中文
相关论文

相关论文: Is human scoring the best criteria for summary eva…

200 篇论文

This paper proposes a general family of estimators for estimating the population mean in systematic sampling in the presence of non-response adapting the family of estimators proposed by Khoshnevisan et al. (2007). In this paper we have…

统计理论 · 数学 2013-05-09 M. K. Chaudhary , Sachin Malik , Jayant Singh , Rajesh Singh

Items in many datasets can be arranged to a natural order. Such orders are useful since they can provide new knowledge about the data and may ease further data exploration and visualization. Our goal in this paper is to define a…

数据结构与算法 · 计算机科学 2019-02-11 Nikolaj Tatti

Evaluating the quality of free-text explanations is a multifaceted, subjective, and labor-intensive task. Large language models (LLMs) present an appealing alternative due to their potential for consistency, scalability, and…

计算与语言 · 计算机科学 2024-09-04 Ana Brassard , Benjamin Heinzerling , Keito Kudo , Keisuke Sakaguchi , Kentaro Inui

We examine evaluation of faithfulness to input data in the context of hotel highlights: brief LLM-generated summaries that capture unique features of accommodations. Through human evaluation campaigns involving categorical error assessment…

计算与语言 · 计算机科学 2025-07-16 Patrícia Schmidtová , Ondřej Dušek , Saad Mahamood

In the presence of weak overall correlation, it may be useful to investigate if the correlation is significantly and substantially more pronounced over a subpopulation. Two different testing procedures are compared. Both are based on the…

机器学习 · 统计学 2015-04-22 Stephen Bamattre , Rex Hu , Joseph S. Verducci

We describe a novel method for efficiently eliciting scalar annotations for dataset construction and system quality estimation by human judgments. We contrast direct assessment (annotators assign scores to items directly), online pairwise…

计算与语言 · 计算机科学 2018-06-05 Keisuke Sakaguchi , Benjamin Van Durme

Knowing when a classifier's prediction can be trusted is useful in many applications and critical for safely using AI. While the bulk of the effort in machine learning research has been towards improving classifier performance,…

机器学习 · 统计学 2018-10-30 Heinrich Jiang , Been Kim , Melody Y. Guan , Maya Gupta

Correlation measure of order $k$ is an important measure of randomness in binary sequences. This measure tries to look for dependence between several shifted version of a sequence. We study the relation between the correlation measure of…

信息论 · 计算机科学 2021-07-27 Zhixiong Chen , Ana I. Gómez , Domingo Gómez-Pérez , Andrew Tirkel

Propensity score plays a central role in causal inference, but its use is not limited to causal comparisons. As a covariate balancing tool, propensity score can be used for controlled descriptive comparisons between groups whose memberships…

统计方法学 · 统计学 2022-09-09 Fan Li , Fan Li

Recommendation systems increasingly depend on massive human-labeled datasets; however, the human annotators hired to generate these labels increasingly come from homogeneous backgrounds. This poses an issue when downstream predictive models…

The direct measurement of quality is difficult because there is no way we can measure quality factors. For measuring these factors, we have to express them in terms of metrics or models. Researchers have developed quality models that…

软件工程 · 计算机科学 2010-04-28 Devpriya Soni , Namita Shrivastava , M. Kumar

Automatic evaluation metrics have been facilitating the rapid development of automatic summarization methods by providing instant and fair assessments of the quality of summaries. Most metrics have been developed for the general domain,…

计算与语言 · 计算机科学 2023-03-21 Hongyi Yuan , Yaoyun Zhang , Fei Huang , Songfang Huang

Large language model (LLM) judges have often been used alongside traditional, algorithm-based metrics for tasks like summarization because they better capture semantic information, are better at reasoning, and are more robust to…

This work offers a novel view on the use of human input as labels, acknowledging that humans may err. We build a behavioral profile for human annotators which is used as a feature representation of the provided input. We show that by…

数据库 · 计算机科学 2022-05-09 Roee Shraga

We address the rating-inference problem, wherein rather than simply decide whether a review is "thumbs up" or "thumbs down", as in previous sentiment analysis work, one must determine an author's evaluation with respect to a multi-point…

计算与语言 · 计算机科学 2007-05-23 Bo Pang , Lillian Lee

Large technology firms face the problem of moderating content on their online platforms for compliance with laws and policies. To accomplish this at the scale of billions of pieces of content per day, a combination of human and machine…

应用统计 · 统计学 2023-06-14 Xuan Yang , Andrew J Smart , Daniel Theron

Field experiments (A/B tests) are often the most credible benchmark for methods (algorithms) in societal systems, but their cost and latency bottleneck rapid methodological progress. LLM-based persona simulation offers a cheap synthetic…

人工智能 · 计算机科学 2026-01-30 Enoch Hyunwook Kang

When customers are faced with the task of making a purchase in an unfamiliar product domain, it might be useful to provide them with an overview of the product set to help them understand what they can expect. In this paper we present and…

计算与语言 · 计算机科学 2018-07-18 Kittipitch Kuptavanich

Human evaluation has been the gold standard for checking faithfulness in abstractive summarization. However, with a challenging source domain like narrative, multiple annotators can agree a summary is faithful, while missing details that…

人工智能 · 计算机科学 2025-04-02 Melanie Subbiah , Faisal Ladhak , Akankshya Mishra , Griffin Adams , Lydia B. Chilton , Kathleen McKeown

This paper investigates the ability of large language models (LLMs) to solve statistical tasks, as well as their capacity to assess the quality of reasoning. While state-of-the-art LLMs have demonstrated remarkable performance in a range of…

计算与语言 · 计算机科学 2026-01-22 Crish Nagarkar , Leonid Bogachev , Serge Sharoff