中文
相关论文

相关论文: Validating LLM-as-a-Judge Systems under Rating Ind…

200 篇论文

The unjudged document problem, where systems that did not contribute to the original judgement pool may retrieve documents without a relevance judgement, is a key obstacle to the reuseability of test collections in information retrieval.…

信息检索 · 计算机科学 2026-01-26 Lukas Gienapp , Martin Potthast , Andrew Yates , Harrisen Scells , Eugene Yang

Large language models (LLMs) can act as evaluators, a role studied by methods like LLM-as-a-Judge and fine-tuned judging LLMs. In the field of education, LLMs have been studied as assistant tools for students and teachers. Our research…

计算与语言 · 计算机科学 2025-09-26 Valeria Ramirez-Garcia , David de-Fitero-Dominguez , Antonio Garcia-Cabot , Eva Garcia-Lopez

Previous work adopts large language models (LLMs) as evaluators to evaluate natural language process (NLP) tasks. However, certain shortcomings, e.g., fairness, scope, and accuracy, persist for current LLM evaluators. To analyze whether…

计算与语言 · 计算机科学 2025-01-22 Qintong Li , Leyang Cui , Lingpeng Kong , Wei Bi

Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating queries without considering user experiences, inadequately…

Large Language Models (LLMs) have demonstrated remarkable progress through preference-based fine-tuning, which critically depends on the quality of the underlying training data. While human feedback is essential for improving data quality,…

人工智能 · 计算机科学 2025-10-31 Derin Cayir , Renjie Tao , Rashi Rungta , Kai Sun , Sean Chen , Haidar Khan , Minseok Kim , Julia Reinspach , Yue Liu

Recently, there has been a growing trend of utilizing Large Language Model (LLM) to evaluate the quality of other LLMs. Many studies have fine-tuned judge models based on open-source LLMs for evaluation. While the fine-tuned judge models…

计算与语言 · 计算机科学 2025-06-02 Hui Huang , Xingyuan Bu , Hongli Zhou , Yingqi Qu , Jing Liu , Muyun Yang , Bing Xu , Tiejun Zhao

Human evaluation for natural language generation (NLG) often suffers from inconsistent user ratings. While previous research tends to attribute this problem to individual user preferences, we show that the quality of human judgements can…

计算与语言 · 计算机科学 2018-10-03 Jekaterina Novikova , Ondřej Dušek , Verena Rieser

Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence. These models not only excel in traditional vision-language tasks but also demonstrate impressive performance in contemporary…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Xiaotian Han , Quanzeng You , Yongfei Liu , Wentao Chen , Huangjie Zheng , Khalil Mrini , Xudong Lin , Yiqi Wang , Bohan Zhai , Jianbo Yuan , Heng Wang , Hongxia Yang

Evaluating image editing models remains challenging due to the coarse granularity and limited interpretability of traditional metrics, which often fail to capture aspects important to human perception and intent. Such metrics frequently…

The growing scale of evaluation tasks has led to the widespread adoption of automated evaluation using LLMs, a paradigm known as "LLM-as-a-judge". However, improving its alignment with human preferences without complex prompts or…

计算与语言 · 计算机科学 2025-10-17 Peng Lai , Jianjie Zheng , Sijie Cheng , Yun Chen , Peng Li , Yang Liu , Guanhua Chen

Large-scale public deliberations generate thousands of free-form contributions that must be synthesized into representative and neutral summaries for policy use. While LLMs have been shown as a promising tool to generate summaries for…

计算与语言 · 计算机科学 2026-03-23 Shenzhe Zhu , Shu Yang , Michiel A. Bakker , Alex Pentland , Jiaxin Pei

Automatic grading of subjective questions remains a significant challenge in examination assessment due to the diversity in question formats and the open-ended nature of student responses. Existing works primarily focus on a specific type…

计算与语言 · 计算机科学 2025-10-10 Fanwei Zhua , Jiaxuan He , Xiaoxiao Chen , Zulong Chen , Quan Lu , Chenrui Mei

Serendipity-oriented recommender systems aim to counteract over-specialization in user preferences. However, evaluating a user's serendipitous response towards a recommended item can be challenging because of its emotional nature. In this…

信息检索 · 计算机科学 2024-12-18 Yu Tokutake , Kazushi Okamoto

Scientific discovery begins with ideas, yet evaluating early-stage research concepts is a subtle and subjective human judgment. As large language models (LLMs) are increasingly tasked with generating scientific hypotheses, most systems…

人机交互 · 计算机科学 2026-03-26 Lingyu Zhang , Mitchell Wang , Boyuan Chen

Evaluating Large Language Models (LLMs) often requires costly human annotations. To address this, LLM-based judges have been proposed, which compare the outputs of two LLMs enabling the ranking of models without human intervention. While…

计算与语言 · 计算机科学 2025-05-28 David Salinas , Omar Swelam , Frank Hutter

Evaluating large language models (LLMs) in long-form, knowledge-grounded role-play dialogues remains challenging. This study compares LLM-generated and human-authored responses in multi-turn professional training simulations through human…

计算与语言 · 计算机科学 2025-10-10 Dongxu Lu , Johan Jeuring , Albert Gatt

Retrieval-Augmented Generation (RAG) has proven its effectiveness in alleviating hallucinations for Large Language Models (LLMs). However, existing automated evaluation metrics cannot fairly evaluate the outputs generated by RAG models…

计算与语言 · 计算机科学 2025-02-27 Shuliang Liu , Xinze Li , Zhenghao Liu , Yukun Yan , Cheng Yang , Zheni Zeng , Zhiyuan Liu , Maosong Sun , Ge Yu

With the broad availability of large language models and their ability to generate vast outputs using varied prompts and configurations, determining the best output for a given task requires an intensive evaluation process, one where…

Ensuring that Large Language Models (LLMs) align with the diverse and evolving human values across different regions and cultures remains a critical challenge in AI ethics. Current alignment approaches often yield superficial conformity…

人工智能 · 计算机科学 2025-11-04 Jiahao Wang , Songkai Xue , Jinghui Li , Xiaozhen Wang

Identifying logical errors in complex, incomplete or even contradictory and overall heterogeneous data like students' experimentation protocols is challenging. Recognizing the limitations of current evaluation methods, we investigate the…

人工智能 · 计算机科学 2024-09-20 Arne Bewersdorff , Kathrin Seßler , Armin Baur , Enkelejda Kasneci , Claudia Nerdel
‹ 上一页 1 8 9 10 下一页 ›