中文
相关论文

相关论文: Chatbot Arena: An Open Platform for Evaluating LLM…

200 篇论文

The integration of Large Language Models (LLMs) into healthcare settings has gained significant attention, particularly for question-answering tasks. Given the high-stakes nature of healthcare, it is essential to ensure that LLM-generated…

Recently, the evaluation of Large Language Models has emerged as a popular area of research. The three crucial questions for LLM evaluation are ``what, where, and how to evaluate''. However, the existing research mainly focuses on the first…

人工智能 · 计算机科学 2023-12-19 Yue Zhang , Ming Zhang , Haipeng Yuan , Shichun Liu , Yongyao Shi , Tao Gui , Qi Zhang , Xuanjing Huang

With the growing adoption and capabilities of vision-language models (VLMs) comes the need for benchmarks that capture authentic user-VLM interactions. In response, we create VisionArena, a dataset of 230K real-world conversations between…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Christopher Chou , Lisa Dunlap , Koki Mashita , Krishna Mandal , Trevor Darrell , Ion Stoica , Joseph E. Gonzalez , Wei-Lin Chiang

Evaluating the abilities of large models and manifesting their gaps are challenging. Current benchmarks adopt either ground-truth-based score-form evaluation on static datasets or indistinct textual chatbot-style human preferences…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Zijian Chen , Lirong Deng , Zhengyu Chen , Kaiwei Zhang , Qi Jia , Yuan Tian , Yucheng Zhu , Guangtao Zhai

The recent success of large language models (LLMs) has shown great potential to develop more powerful conversational recommender systems (CRSs), which rely on natural language conversations to satisfy user needs. In this paper, we embark on…

计算与语言 · 计算机科学 2024-06-21 Xiaolei Wang , Xinyu Tang , Wayne Xin Zhao , Jingyuan Wang , Ji-Rong Wen

The integration of Large Language Models (LLMs) and chatbots introduces new challenges and opportunities for decision-making in software testing. Decision-making relies on a variety of information, including code, requirements…

软件工程 · 计算机科学 2024-06-18 Francisco Gomes de Oliveira Neto

Open community-driven platforms like Chatbot Arena that collect user preference data from site visitors have gained a reputation as one of the most trustworthy publicly available benchmarks for LLM performance. While now standard, it is…

人机交互 · 计算机科学 2024-12-06 Wenting Zhao , Alexander M. Rush , Tanya Goyal

With the rapid adoption of LLM-based chatbots, there is a pressing need to evaluate what humans and LLMs can achieve together. However, standard benchmarks, such as MMLU, measure LLM capabilities in isolation (i.e., "AI-alone"). Here, we…

计算与语言 · 计算机科学 2025-08-13 Serina Chang , Ashton Anderson , Jake M. Hofman

Chatbot Arena is a popular platform for evaluating LLMs by pairwise battles, where users vote for their preferred response from two randomly sampled anonymous models. While Chatbot Arena is widely regarded as a reliable LLM ranking…

计算与语言 · 计算机科学 2025-08-12 Rui Min , Tianyu Pang , Chao Du , Qian Liu , Minhao Cheng , Min Lin

The rapid evolution of Large Language Models (LLMs) has outpaced the development of model evaluation, highlighting the need for continuous curation of new, challenging benchmarks. However, manual curation of high-quality, human-aligned…

机器学习 · 计算机科学 2024-10-16 Tianle Li , Wei-Lin Chiang , Evan Frick , Lisa Dunlap , Tianhao Wu , Banghua Zhu , Joseph E. Gonzalez , Ion Stoica

Since its release in November 2022, ChatGPT has shaken up Stack Overflow, the premier platform for developers queries on programming and software development. Demonstrating an ability to generate instant, human-like responses to technical…

软件工程 · 计算机科学 2025-07-10 Leuson Da Silva , Jordan Samhi , Foutse Khomh

We explore how large language models (LLMs) can enhance the proposal selection process at large user facilities, offering a scalable, consistent, and cost-effective alternative to traditional human review. Proposal selection depends on…

人工智能 · 计算机科学 2025-12-12 Lijie Ding , Janell Thomson , Jon Taylor , Changwoo Do

Large language models (LLMs) chatbots like ChatGPT are increasingly used for mental health support. They offer accessible, therapeutic support but also raise concerns about misinformation, over-reliance, and risks in high-stakes contexts of…

计算机与社会 · 计算机科学 2025-12-09 Lingyao Li , Xiaoshan Huang , Renkai Ma , Ben Zefeng Zhang , Haolun Wu , Fan Yang , Chen Chen

This paper investigates the voting behaviors of Large Language Models (LLMs), specifically GPT-4 and LLaMA-2, their biases, and how they align with human voting patterns. Our methodology involved using a dataset from a human voting…

计算与语言 · 计算机科学 2024-12-20 Joshua C. Yang , Damian Dailisan , Marcin Korecki , Carina I. Hausladen , Dirk Helbing

The advent of AI driven large language models (LLMs) have stirred discussions about their role in qualitative research. Some view these as tools to enrich human understanding, while others perceive them as threats to the core values of the…

软件工程 · 计算机科学 2023-06-26 Muneera Bano , Didar Zowghi , Jon Whittle

We propose a method for evaluating the robustness of widely used LLM ranking systems -- variants of a Bradley--Terry model -- to dropping a worst-case very small fraction of preference data. Our approach is computationally fast and easy to…

机器学习 · 统计学 2026-03-06 Jenny Y. Huang , Yunyi Shen , Dennis Wei , Tamara Broderick

The rise of Large Language Models (LLMs), such as LLaMA and ChatGPT, has opened new opportunities for enhancing recommender systems through improved explainability. This paper provides a systematic literature review focused on leveraging…

信息检索 · 计算机科学 2025-01-22 Alan Said

Providing scaffolding through educational chatbots built on Large Language Models (LLM) has potential risks and benefits that remain an open area of research. When students navigate impasses, they ask for help by formulating impasse-driven…

人机交互 · 计算机科学 2026-02-23 Alexandra Neagu , Marcus Messer , Peter Johnson , Rhodri Nelson

The integration of Large Language Models (LLMs) into the healthcare domain has the potential to significantly enhance patient care and support through the development of empathetic, patient-facing chatbots. This study investigates an…

计算与语言 · 计算机科学 2024-05-28 Man Luo , Christopher J. Warren , Lu Cheng , Haidar M. Abdul-Muhsin , Imon Banerjee

When asked, large language models (LLMs) like ChatGPT claim that they can assist with relevance judgments but it is not clear whether automated judgments can reliably be used in evaluations of retrieval systems. In this perspectives paper,…