English
Related papers

Related papers: Generation Challenges: Results of the Accuracy Eva…

200 papers

We introduce a novel task consisting in assigning a proof to a given mathematical statement. The task is designed to improve the processing of research-level mathematical texts. Applying Natural Language Processing (NLP) tools to research…

Computation and Language · Computer Science 2021-02-04 Maximin Coavoux , Shay B. Cohen

Retrieval-Augmented Generation (RAG) has advanced significantly in recent years. The complexity of RAG systems, which involve multiple components-such as indexing, retrieval, and generation-along with numerous other parameters, poses…

Information Retrieval · Computer Science 2025-08-08 Lorenz Brehme , Thomas Ströhle , Ruth Breu

Despite demonstrating remarkable performance across a wide range of tasks, large language models (LLMs) have also been found to frequently produce outputs that are incomplete or selectively omit key information. In sensitive domains, such…

Computation and Language · Computer Science 2026-05-11 Adam Dejl , James Barry , Alessandra Pascale , Javier Carnerero Cano

Evaluating the quality of generated text is difficult, since traditional NLG evaluation metrics, focusing more on surface form than meaning, often fail to assign appropriate scores. This is especially problematic for AMR-to-text evaluation,…

Computation and Language · Computer Science 2022-05-25 Laura Zeidler , Juri Opitz , Anette Frank

Fake news, misinformation, and unverifiable facts on social media platforms propagate disharmony and affect society, especially when dealing with an epidemic like COVID-19. The task of Fake News Detection aims to tackle the effects of such…

Computation and Language · Computer Science 2021-12-14 Mrinal Rawat , Diptesh Kanojia

This study reports the second shared task named as UrduFake@FIRE2021 on identifying fake news detection in Urdu language. This is a binary classification problem in which the task is to classify a given news article into two classes: (i)…

Computation and Language · Computer Science 2022-07-13 Maaz Amjad , Sabur Butt , Hamza Imam Amjad , Grigori Sidorov , Alisa Zhila , Alexander Gelbukh

This paper offers a comprehensive review of the research on Natural Language Generation (NLG) over the past two decades, especially in relation to data-to-text generation and text-to-text generation deep learning methods, as well as new…

Computation and Language · Computer Science 2022-08-03 Chenhe Dong , Yinghui Li , Haifan Gong , Miaoxin Chen , Junxin Li , Ying Shen , Min Yang

Evaluating natural language generation (NLG) is a vital but challenging problem in natural language processing. Traditional evaluation metrics mainly capturing content (e.g. n-gram) overlap between system outputs and references are far from…

Computation and Language · Computer Science 2025-05-15 Mingqi Gao , Xinyu Hu , Jie Ruan , Xiao Pu , Xiaojun Wan

Fact-checking is the process of evaluating the veracity of claims (i.e., purported facts). In this opinion piece, we raise an issue that has received little attention in prior work -- that some claims are far more difficult to fact-check…

Computation and Language · Computer Science 2022-02-08 Prakhar Singh , Anubrata Das , Junyi Jessy Li , Matthew Lease

We present SemEval-2019 Task 8 on Fact Checking in Community Question Answering Forums, which features two subtasks. Subtask A is about deciding whether a question asks for factual information vs. an opinion/advice vs. just socializing.…

Computation and Language · Computer Science 2019-06-06 Tsvetomila Mihaylova , Georgi Karadjov , Pepa Atanasova , Ramy Baly , Mitra Mohtarami , Preslav Nakov

Evaluating LLM-generated text has become a key challenge, especially in domain-specific contexts like the medical field. This work introduces a novel evaluation methodology for LLM-generated medical explanatory arguments, relying on Proxy…

Computation and Language · Computer Science 2024-10-01 Iker De la Iglesia , Iakes Goenaga , Johanna Ramirez-Romero , Jose Maria Villa-Gonzalez , Josu Goikoetxea , Ander Barrena

We present an overview of the FIGNEWS shared task, organized as part of the ArabicNLP 2024 conference co-located with ACL 2024. The shared task addresses bias and propaganda annotation in multilingual news posts. We focus on the early days…

Computation and Language · Computer Science 2024-07-26 Wajdi Zaghouani , Mustafa Jarrar , Nizar Habash , Houda Bouamor , Imed Zitouni , Mona Diab , Samhaa R. El-Beltagy , Muhammed AbuOdeh

We present the results from the second shared task on multimodal machine translation and multilingual image description. Nine teams submitted 19 systems to two tasks. The multimodal translation task, in which the source sentence is…

Computation and Language · Computer Science 2017-10-20 Desmond Elliott , Stella Frank , Loïc Barrault , Fethi Bougares , Lucia Specia

Machine learning approaches applied to NLP are often evaluated by summarizing their performance in a single number, for example accuracy. Since most test sets are constructed as an i.i.d. sample from the overall data, this approach overly…

Question semantic similarity (Q2Q) is a challenging task that is very useful in many NLP applications, such as detecting duplicate questions and question answering systems. In this paper, we present the results and findings of the shared…

Computation and Language · Computer Science 2019-09-24 Haitham Seelawi , Ahmad Mustafa , Hesham Al-Bataineh , Wael Farhan , Hussein T. Al-Natsheh

In this paper, we present a system that uses a Large Language Model (LLM) to perform grammar and spelling correction as a component of Quality Assurance (QA) for texts generated by NLG systems, which is important for text production in…

Computation and Language · Computer Science 2025-01-28 Ching-Yi Chen , Johanna Heininger , Adela Schneider , Christian Eckard , Andreas Madsack , Robert Weißgraeber

This work discusses an important issue in the area of human resource management by proposing a novel model for creation and evaluation of software teams. The model consists of several assessments, including a technical test, a quality of…

Computers and Society · Computer Science 2015-12-03 Surayne Torres , Yadenis Pinero , Pedro Pinero , Luiz Fernando Capretz

Social networking sites, blogs, and online articles are instant sources of news for internet users globally. However, in the absence of strict regulations mandating the genuineness of every text on social media, it is probable that some of…

Computation and Language · Computer Science 2022-12-08 Arjun Choudhry , Inder Khatri , Minni Jain , Dinesh Kumar Vishwakarma

Collecting human judgements is currently the most reliable evaluation method for natural language generation systems. Automatic metrics have reported flaws when applied to measure quality aspects of generated text and have been shown to…

Computation and Language · Computer Science 2022-04-29 Thórhildur Thorleiksdóttir , Cedric Renggli , Nora Hollenstein , Ce Zhang

Automatic evaluation of generative tasks using large language models faces challenges due to ambiguous criteria. Although automatic checklist generation is a potentially promising approach, its usefulness remains underexplored. We…

Computation and Language · Computer Science 2025-08-22 Momoka Furuhashi , Kouta Nakayama , Takashi Kodama , Saku Sugawara
‹ Prev 1 3 4 5 6 7 10 Next ›