中文
相关论文

相关论文: Towards Personalized Evaluation of Large Language …

200 篇论文

This study explores the use of Large Language Models (LLMs) to analyze text comments from Reddit users, aiming to achieve two primary objectives: firstly, to pinpoint critical excerpts that support a predefined psychological assessment of…

计算与语言 · 计算机科学 2024-02-07 Sergi Blanco-Cuaresma

The unprecedented performance of large language models (LLMs) requires comprehensive and accurate evaluation. We argue that for LLMs evaluation, benchmarks need to be comprehensive and systematic. To this end, we propose the ZhuJiu…

计算与语言 · 计算机科学 2023-08-29 Baoli Zhang , Haining Xie , Pengfan Du , Junhao Chen , Pengfei Cao , Yubo Chen , Shengping Liu , Kang Liu , Jun Zhao

Crowdsourcing is an easy, cheap, and fast way to perform large scale quality assessment; however, human judgments are often influenced by cognitive biases, which lowers their credibility. In this study, we focus on cognitive biases…

人机交互 · 计算机科学 2024-07-30 Shun Ito , Hisashi Kashima

Large Language Models (LLMs) are rapidly evolving and impacting various fields, necessitating the development of effective methods to evaluate and compare their performance. Most current approaches for performance evaluation are either…

计算与语言 · 计算机科学 2025-02-11 Behrad Moniri , Hamed Hassani , Edgar Dobriban

Large Language Models (LLMs) have unlocked new capabilities and applications; however, evaluating the alignment with human preferences still poses significant challenges. To address this issue, we introduce Chatbot Arena, an open platform…

Evaluating large language models typically relies on human-authored benchmarks, reference answers, and human or single-model judgments, approaches that scale poorly, become quickly outdated, and mismatch open-world deployments that depend…

人工智能 · 计算机科学 2026-02-04 Yanki Margalit , Erni Avram , Ran Taig , Oded Margalit , Nurit Cohen-Inger

Generative Machine Learning models have become central to modern systems, powering applications in creative writing, summarization, multi-hop reasoning, and context-aware dialogue. These models underpin large-scale AI assistants, workflow…

机器学习 · 计算机科学 2025-08-08 Arthur Cho

Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have ushered in a new era of AI capabilities, demonstrating near-human-level performance across diverse scenarios. While numerous benchmarks (e.g., MMLU) and…

人工智能 · 计算机科学 2025-09-03 Kangyu Wang , Hongliang He , Lin Liu , Ruiqi Liang , Zhenzhong Lan , Jianguo Li

Recently, the evaluation of Large Language Models has emerged as a popular area of research. The three crucial questions for LLM evaluation are ``what, where, and how to evaluate''. However, the existing research mainly focuses on the first…

人工智能 · 计算机科学 2023-12-19 Yue Zhang , Ming Zhang , Haipeng Yuan , Shichun Liu , Yongyao Shi , Tao Gui , Qi Zhang , Xuanjing Huang

Online questionnaires that use crowd-sourcing platforms to recruit participants have become commonplace, due to their ease of use and low costs. Artificial Intelligence (AI) based Large Language Models (LLM) have made it easy for bad actors…

人机交互 · 计算机科学 2024-02-02 Benjamin Lebrun , Sharon Temtsin , Andrew Vonasch , Christoph Bartneck

The advent of large language models marks a revolutionary breakthrough in artificial intelligence. With the unprecedented scale of training and model parameters, the capability of large language models has been dramatically improved,…

Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current benchmarks face three critical limitations: (1) fragmented evaluation dimensions that often…

Overestimation in evaluating large language models (LLMs) has become an increasing concern. Due to the contamination of public benchmarks or imbalanced model training, LLMs may achieve unreal evaluation results on public benchmarks, either…

计算与语言 · 计算机科学 2026-05-26 Zi Liang , Liantong Yu , Shiyu Zhang , Qingqing Ye , Haibo Hu

Multimodal Large Language Models (MLLMs) have become increasingly important due to their state-of-the-art performance and ability to integrate multiple data modalities, such as text, images, and audio, to perform complex tasks with high…

Autonomous agents have long been a prominent research focus in both academic and industry communities. Previous research in this field often focuses on training agents with limited knowledge within isolated environments, which diverges…

Recently, various Large Language Models (LLMs) evaluation datasets have emerged, but most of them have issues with distorted rankings and difficulty in model capabilities analysis. Addressing these concerns, this paper introduces ANGO, a…

计算与语言 · 计算机科学 2024-02-22 Bingchao Wang

A recent focus of large language model (LLM) development, as exemplified by generative search engines, is to incorporate external references to generate and support its claims. However, evaluating the attribution, i.e., verifying whether…

计算与语言 · 计算机科学 2023-10-10 Xiang Yue , Boshi Wang , Ziru Chen , Kai Zhang , Yu Su , Huan Sun

Online reviews play a pivotal role in influencing consumer decisions across various domains, from purchasing products to selecting hotels or restaurants. However, the sheer volume of reviews -- often containing repetitive or irrelevant…

计算与语言 · 计算机科学 2024-10-18 Mir Tafseer Nayeem , Davood Rafiei

In this paper, we present empirical studies on conversational recommendation tasks using representative large language models in a zero-shot setting with three primary contributions. (1) Data: To gain insights into model behavior in…

‹ 上一页 1 2 3 10 下一页 ›