English
Related papers

Related papers: UCSC at SemEval-2025 Task 3: Context, Models and P…

200 papers

We present the Mu-SHROOM shared task which is focused on detecting hallucinations and other overgeneration mistakes in the output of instruction-tuned large language models (LLMs). Mu-SHROOM addresses general-purpose LLMs in 14 languages,…

This paper describes our submission for SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared-task on Hallucinations and Related Observable Overgeneration Mistakes. The task involves detecting hallucinated spans in text generated by…

Computation and Language · Computer Science 2025-05-28 Baraa Hikal , Ahmed Nasreldin , Ali Hamdi

Hallucinations are one of the major problems of LLMs, hindering their trustworthiness and deployment to wider use cases. However, most of the research on hallucinations focuses on English data, neglecting the multilingual nature of LLMs.…

Computation and Language · Computer Science 2025-07-02 Miriam Anschütz , Ekaterina Gikalo , Niklas Herbster , Georg Groh

Identification of hallucination spans in black-box language model generated text is essential for applications in the real world. A recent attempt at this direction is SemEval-2025 Task 3, Mu-SHROOM-a Multilingual Shared Task on…

Computation and Language · Computer Science 2025-05-26 Saketh Reddy Vemula , Parameswari Krishnamurthy

This paper presents our findings of the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes, MU-SHROOM, which focuses on identifying hallucinations and related overgeneration errors in large language…

SemEval-2025 Task 3 (Mu-SHROOM) focuses on detecting hallucinations in content generated by various large language models (LLMs) across multiple languages. This task involves not only identifying the presence of hallucinations but also…

Computation and Language · Computer Science 2025-05-13 Jiaying Hong , Thanet Markchom , Jianfei Xu , Tong Wu , Huizhi Liang

We present the system developed by the Central China Normal University (CCNU) team for the Mu-SHROOM shared task, which focuses on identifying hallucinations in question-answering systems across 14 different languages. Our approach…

Computation and Language · Computer Science 2025-05-20 Xu Liu , Guanyi Chen

Multilingual hallucination detection stands as an underexplored challenge, which the Mu-SHROOM shared task seeks to address. In this work, we propose an efficient, training-free LLM prompting strategy that enhances detection by translating…

Computation and Language · Computer Science 2025-08-04 Dimitra Karkani , Maria Lymperaiou , Giorgos Filandrianos , Nikolaos Spanos , Athanasios Voulodimos , Giorgos Stamou

This paper presents the contributions of the ATLANTIS team to SemEval-2025 Task 3, focusing on detecting hallucinated text spans in question answering systems. Large Language Models (LLMs) have significantly advanced Natural Language…

Computation and Language · Computer Science 2025-08-08 Catherine Kobus , François Lancelot , Marion-Cécile Martin , Nawal Ould Amer

Hallucinations in large language models (LLMs) have recently become a significant problem. A recent effort in this direction is a shared task at Semeval 2024 Task 6, SHROOM, a Shared-task on Hallucinations and Related Observable…

Computation and Language · Computer Science 2024-04-12 Rahul Mehta , Andrew Hoblitzell , Jack O'Keefe , Hyeju Jang , Vasudeva Varma

This paper presents the results of the SHROOM, a shared task focused on detecting hallucinations: outputs from natural language generation (NLG) systems that are fluent, yet inaccurate. Such cases of overgeneration put in jeopardy many NLG…

In this paper, we present our team's submissions for SemEval-2024 Task-6 - SHROOM, a Shared-task on Hallucinations and Related Observable Overgeneration Mistakes. The participants were asked to perform binary classification to identify…

Computation and Language · Computer Science 2024-04-15 Natalia Grigoriadou , Maria Lymperaiou , Giorgos Filandrianos , Giorgos Stamou

In this paper, we present HalluSearch, a multilingual pipeline designed to detect fabricated text spans in Large Language Model (LLM) outputs. Developed as part of Mu-SHROOM, the Multilingual Shared-task on Hallucinations and Related…

Computation and Language · Computer Science 2025-04-15 Mohamed A. Abdallah , Samhaa R. El-Beltagy

Language models, particularly generative models, are susceptible to hallucinations, generating outputs that contradict factual knowledge or the source text. This study explores methods for detecting hallucinations in three SemEval-2024 Task…

Detecting spans of hallucination in LLM-generated answers is crucial for improving factual consistency. This paper presents a span-level hallucination detection framework for the SemEval-2025 Shared Task, focusing on English and Arabic…

Computation and Language · Computer Science 2025-04-29 Passant Elchafei , Mervet Abu-Elkheir

We describe the University of Amsterdam Intelligent Data Engineering Lab team's entry for the SemEval-2024 Task 6 competition. The SHROOM-INDElab system builds on previous work on using prompt programming and in-context learning with large…

Computation and Language · Computer Science 2024-04-08 Bradley P. Allen , Fina Polat , Paul Groth

In Natural Language Generation (NLG), contemporary Large Language Models (LLMs) face several challenges, such as generating fluent yet inaccurate outputs and reliance on fluency-centric metrics. This often leads to neural networks…

Computation and Language · Computer Science 2025-12-19 Federico Borra , Claudio Savelli , Giacomo Rosso , Alkis Koudounas , Flavio Giobergia

Hallucinations in large language model (LLM) outputs severely limit their reliability in knowledge-intensive tasks such as question answering. To address this challenge, we introduce REFIND (Retrieval-augmented Factuality hallucINation…

Computation and Language · Computer Science 2025-04-09 DongGeon Lee , Hwanjo Yu

This paper mainly describes a unified system for hallucination detection of LLMs, which wins the second prize in the model-agnostic track of the SemEval-2024 Task 6, and also achieves considerable results in the model-aware track. This task…

Computation and Language · Computer Science 2024-02-21 Chengcheng Wei , Ze Chen , Songtan Fang , Jiarong He , Max Gao

Large language models (LLMs) have achieved impressive performance across a wide range of natural language processing tasks, yet they often produce hallucinated content that undermines factual reliability. To address this challenge, we…

Computation and Language · Computer Science 2026-03-23 Yaxin Zhao , Yu Zhang
‹ Prev 1 2 3 10 Next ›