English
Related papers

Related papers: Chatbot Arena Meets Nuggets: Towards Explanations …

200 papers

Large Language Models (LLMs) have significantly enhanced the capabilities of information access systems, especially with retrieval-augmented generation (RAG). Nevertheless, the evaluation of RAG systems remains a barrier to continued…

Information Retrieval · Computer Science 2025-04-22 Ronak Pradeep , Nandan Thakur , Shivani Upadhyay , Daniel Campos , Nick Craswell , Jimmy Lin

As LLMs continuously evolve, there is an urgent need for a reliable evaluation method that delivers trustworthy results promptly. Currently, static benchmarks suffer from inflexibility and unreliability, leading users to prefer human voting…

Computation and Language · Computer Science 2024-10-08 Ruochen Zhao , Wenxuan Zhang , Yew Ken Chia , Weiwen Xu , Deli Zhao , Lidong Bing

Evaluation of long-form, citation-backed reports has lately received significant attention due to the wide-scale adoption of retrieval-augmented generation (RAG) systems. Core to many evaluation frameworks is the use of atomic facts, or…

Computation and Language · Computer Science 2026-05-07 Bryan Li , William Walden , Yu Hou , Gabrielle Kaili-May Liu , Dawn Lawrie , Jame Mayfield , Eugene Yang , Chris Callison-Burch , Laura Dietz

This paper examines the application of ChatGPT, a large language model (LLM), for question-and-answer (Q&A) tasks in the highly specialized field of nuclear data. The primary focus is on evaluating ChatGPT's performance on a curated test…

Computation and Language · Computer Science 2024-09-04 Muhammad Anwar , Mischa de Costa , Issam Hammad , Daniel Lau

The rise of personalized conversational search systems has been driven by advancements in Large Language Models (LLMs), enabling these systems to retrieve and generate answers for complex information needs. However, the automatic evaluation…

Information Retrieval · Computer Science 2025-03-14 Zahra Abbasiantaeb , Simon Lupart , Leif Azzopardi , Jeffery Dalton , Mohammad Aliannejadi

Large Language Models (LLMs) have unlocked new capabilities and applications; however, evaluating the alignment with human preferences still poses significant challenges. To address this issue, we introduce Chatbot Arena, an open platform…

Evaluating the reasoning abilities of large language models (LLMs) is challenging. Existing benchmarks often depend on static datasets, which are vulnerable to data contamination and may get saturated over time, or on binary live human…

Artificial Intelligence · Computer Science 2025-02-18 Lanxiang Hu , Qiyu Li , Anze Xie , Nan Jiang , Ion Stoica , Haojian Jin , Hao Zhang

Unlike short-form retrieval-augmented generation (RAG), such as factoid question answering, long-form RAG requires retrieval to provide documents covering a wide range of relevant information. Automated report generation exemplifies this…

Information Retrieval · Computer Science 2026-01-30 Jia-Huei Ju , François G. Landry , Eugene Yang , Suzan Verberne , Andrew Yates

Evaluating the quality of retrieval-augmented generation (RAG) and document reranking systems remains challenging due to the lack of scalable, user-centric, and multi-perspective evaluation tools. We introduce RankArena, a unified platform…

Information Retrieval · Computer Science 2025-08-08 Abdelrahman Abdallah , Mahmoud Abdalla , Bhawna Piryani , Jamshid Mozafari , Mohammed Ali , Adam Jatowt

Large Language Models (LLMs) have proven immensely beneficial in education by capturing vast amounts of literature-based information, allowing them to generate context without relying on external sources. In this paper, we propose a…

Information Retrieval · Computer Science 2025-07-03 Umar Ali Khan , Ekram Khan , Fiza Khan , Athar Ali Moinuddin

While large language models (LLMs) are increasingly used to summarize long documents, this trend poses significant challenges in the legal domain, where the factual accuracy of deposition summaries is crucial. Nugget-based methods have been…

Computation and Language · Computer Science 2026-01-22 Naghmeh Farzi , Laura Dietz , Dave D. Lewis

Question answering systems (QA) utilizing Large Language Models (LLMs) heavily depend on the retrieval component to provide them with domain-specific information and reduce the risk of generating inaccurate responses or hallucinations.…

Computation and Language · Computer Science 2024-06-11 Ashkan Alinejad , Krtin Kumar , Ali Vahdat

RAG systems are increasingly evaluated and optimized using LLM judges, an approach that is rapidly becoming the dominant paradigm for system assessment. Nugget-based approaches in particular are now embedded not only in evaluation…

Information Retrieval · Computer Science 2026-03-30 Laura Dietz , Bryan Li , Eugene Yang , Dawn Lawrie , William Walden , James Mayfield

Evaluating the quality of recommender systems is critical for algorithm design and optimization. Most evaluation methods are computed based on offline metrics for quick algorithm evolution, since online experiments are usually risky and…

Information Retrieval · Computer Science 2024-12-17 Zhuo Wu , Qinglin Jia , Chuhan Wu , Zhaocheng Du , Shuai Wang , Zan Wang , Zhenhua Dong

This report provides an initial look at partial results from the TREC 2024 Retrieval-Augmented Generation (RAG) Track. We have identified RAG evaluation as a barrier to continued progress in information access (and more broadly, natural…

Information Retrieval · Computer Science 2024-11-15 Ronak Pradeep , Nandan Thakur , Shivani Upadhyay , Daniel Campos , Nick Craswell , Jimmy Lin

Assessing the effectiveness of large language models (LLMs) presents substantial challenges. The method of conducting human-annotated battles in an online Chatbot Arena is a highly effective evaluative technique. However, this approach is…

Computation and Language · Computer Science 2024-07-16 Haipeng Luo , Qingfeng Sun , Can Xu , Pu Zhao , Qingwei Lin , Jianguang Lou , Shifeng Chen , Yansong Tang , Weizhu Chen

Search-augmented language models combine web search with Large Language Models (LLMs) to improve response groundedness and freshness. However, analyzing these systems remains challenging: existing datasets are limited in scale and narrow in…

Computational argumentation, which involves generating answers or summaries for controversial topics like abortion bans and vaccination, has become increasingly important in today's polarized environment. Sophisticated LLM capabilities…

Computation and Language · Computer Science 2024-12-09 Kaustubh D. Dhole , Kai Shu , Eugene Agichtein

Retrieval-augmented generation (RAG) is a popular technique for using large language models (LLMs) to build customer-support, question-answering solutions. In this paper, we share our team's practical experience building and maintaining…

Information Retrieval · Computer Science 2024-10-18 Sarah Packowski , Inge Halilovic , Jenifer Schlotfeldt , Trish Smith

Question answering based on retrieval augmented generation (RAG-QA) is an important research topic in NLP and has a wide range of real-world applications. However, most existing datasets for this task are either constructed using a single…

Computation and Language · Computer Science 2024-10-04 Rujun Han , Yuhao Zhang , Peng Qi , Yumo Xu , Jenyuan Wang , Lan Liu , William Yang Wang , Bonan Min , Vittorio Castelli
‹ Prev 1 2 3 10 Next ›