中文
相关论文

相关论文: Sensitive and Scalable Online Evaluation with Theo…

200 篇论文

Online ranker evaluation is a key challenge in information retrieval. An important task in the online evaluation of rankers is using implicit user feedback for inferring preferences between rankers. Interleaving methods have been found to…

信息检索 · 计算机科学 2016-08-03 Brian Brost , Ingemar J. Cox , Yevgeny Seldin , Christina Lioma

As large language models (LLMs) are increasingly used as evaluators for natural language generation tasks, ensuring unbiased assessments is essential. However, LLM evaluators often display biased preferences, such as favoring verbosity and…

计算与语言 · 计算机科学 2025-04-21 Hawon Jeong , ChaeHun Park , Jimin Hong , Hojoon Lee , Jaegul Choo

Large Language Models (LLMs) have demonstrated promising capabilities as automatic evaluators in assessing the quality of generated natural language. However, LLMs still exhibit biases in evaluation and often struggle to generate coherent…

计算与语言 · 计算机科学 2025-01-20 Yinhong Liu , Han Zhou , Zhijiang Guo , Ehsan Shareghi , Ivan Vulić , Anna Korhonen , Nigel Collier

The advent of large language models (LLMs) offers unprecedented opportunities to reimagine peer review beyond the constraints of traditional workflows. Despite these opportunities, prior efforts have largely focused on replicating…

计算与语言 · 计算机科学 2025-09-26 Yaohui Zhang , Haijing Zhang , Wenlong Ji , Tianyu Hua , Nick Haber , Hancheng Cao , Weixin Liang

With the onset of large language models (LLMs), the performance of artificial intelligence (AI) models is becoming increasingly multi-dimensional. Accordingly, there have been several large, multi-dimensional evaluation frameworks put…

人机交互 · 计算机科学 2025-06-05 Sean Steinle

Personalization plays an important role in many services. To evaluate personalized rankings, online evaluation, such as A/B testing, is widely used today. Recently, multileaving has been found to be an efficient method for evaluating…

信息检索 · 计算机科学 2019-07-22 Kojiro Iizuka , Takeshi Yoneda , Yoshifumi Seki

This study presents a theoretical analysis on the efficiency of interleaving, an efficient online evaluation method for rankings. Although interleaving has already been applied to production systems, the source of its high efficiency has…

信息检索 · 计算机科学 2023-06-21 Kojiro Iizuka , Hajime Morita , Makoto P. Kato

News recommendation is a challenging task that involves personalization based on the interaction history and preferences of each user. Recent works have leveraged the power of pretrained language models (PLMs) to directly rank news items by…

信息检索 · 计算机科学 2024-09-27 Nithish Kannen , Yao Ma , Gerrit J. J. van den Burg , Jean Baptiste Faddoul

Interleaving is an online evaluation approach for information retrieval systems that compares the effectiveness of ranking functions in interpreting the users' implicit feedback. Previous work such as Hofmann et al (2011) has evaluated the…

信息检索 · 计算机科学 2023-03-20 Alessandro Benedetti , Anna Ruggero

Most popular strategies to capture subjective judgments from humans involve the construction of a unidimensional relative measurement scale, representing order preferences or judgments about a set of objects or conditions. This information…

应用统计 · 统计学 2017-12-18 Maria Perez-Ortiz , Rafal K. Mantiuk

Large language models (LLMs) have shown remarkable success, but aligning them with human preferences remains a core challenge. As individuals have their own, multi-dimensional preferences, recent studies have explored multi-dimensional…

机器学习 · 计算机科学 2025-06-03 Minhyeon Oh , Seungjoon Lee , Jungseul Ok

Subjective assessment tests are often employed to evaluate image processing systems, notably image and video compression, super-resolution among others and have been used as an indisputable way to provide evidence of the performance of an…

多媒体 · 计算机科学 2023-11-13 Shima Mohammadi , Joao Ascenso

Human evaluation of generated language through pairwise preference judgments is pervasive. However, under common scenarios, such as when generations from a model pair are very similar, or when stochastic decoding results in large variations…

计算与语言 · 计算机科学 2024-10-30 Sayan Ghosh , Tejas Srinivasan , Swabha Swayamdipta

The adoption of Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks. We identify two fundamental types of inconsistencies: (1) Score-Comparison…

Deciding which large language model (LLM) to use is a complex challenge. Pairwise ranking has emerged as a new method for evaluating human preferences for LLMs. This approach entails humans evaluating pairs of model outputs based on a…

计算与语言 · 计算机科学 2025-02-18 Roland Daynauth , Christopher Clarke , Krisztian Flautner , Lingjia Tang , Jason Mars

Recent advances have made long-form report-generating systems widely available. This has prompted evaluation frameworks that use LLM-as-judge protocols and claim verification, along with meta-evaluation frameworks that seek to validate…

Despite the retrieval effectiveness of queries being mutually independent of one another, the evaluation of query performance prediction (QPP) systems has been carried out by measuring rank correlation over an entire set of queries. Such a…

信息检索 · 计算机科学 2023-04-04 Suchana Datta , Debasis Ganguly , Derek Greene , Mandar Mitra

Evaluating the causal effect of recommendations is an important objective because the causal effect on user interactions can directly leads to an increase in sales and user engagement. To select an optimal recommendation model, it is common…

机器学习 · 计算机科学 2021-07-16 Masahiro Sato

Rating-based human evaluation has become an essential tool to accurately evaluate the impressive performance of large language models (LLMs). However, current rating systems suffer from several important limitations: first, they fail to…

计算与语言 · 计算机科学 2025-02-12 Jasper Dekoninck , Maximilian Baader , Martin Vechev

Assessing image quality is crucial in image processing tasks such as compression, super-resolution, and denoising. While subjective assessments involving human evaluators provide the most accurate quality scores, they are impractical for…

多媒体 · 计算机科学 2025-03-26 Shima Mohammadi , João Ascenso
‹ 上一页 1 2 3 10 下一页 ›