English
Related papers

Related papers: UPB at IberLEF-2023 AuTexTification: Detection of …

200 papers

The Multigenerator, Multidomain, and Multilingual Black-Box Machine-Generated Text Detection shared task in the SemEval-2024 competition aims to tackle the problem of misusing collaborative human-AI writing. Although there are a lot of…

Computation and Language · Computer Science 2024-05-20 Anastasia Voznyuk , Vasily Konovalov

Most Reading Comprehension methods limit themselves to queries which can be answered using a single sentence, paragraph, or document. Enabling models to combine disjoint pieces of textual evidence would extend the scope of machine…

Computation and Language · Computer Science 2018-06-12 Johannes Welbl , Pontus Stenetorp , Sebastian Riedel

Large Language Models (LLMs) have achieved human-level fluency in text generation, making it difficult to distinguish between human-written and LLM-generated texts. This poses a growing risk of misuse of LLMs and demands the development of…

Computation and Language · Computer Science 2024-02-20 Ryuto Koike , Masahiro Kaneko , Naoaki Okazaki

The pre-training of large language models usually requires massive amounts of resources, both in terms of computation and data. Frequently used web sources such as Common Crawl might contain enough noise to make this pre-training…

Computation and Language · Computer Science 2022-07-15 Javier de la Rosa , Eduardo G. Ponferrada , Paulo Villegas , Pablo Gonzalez de Prado Salas , Manu Romero , Marıa Grandury

Large semantic knowledge bases are grounded in factual knowledge. However, recent approaches to dense text representations (i.e. embeddings) do not efficiently exploit these resources. Dense and robust representations of documents are…

Artificial Intelligence · Computer Science 2024-10-01 Boshko Koloski , Senja Pollak , Roberto Navigli , Blaž Škrlj

Entity coreference resolution is an important research problem with many applications, including information extraction and question answering. Coreference resolution for English has been studied extensively. However, there is relatively…

Computation and Language · Computer Science 2023-01-24 Tuan Manh Lai , Heng Ji

Machine learning models for text classification are trained to predict a class for a given text. To do this, training and validation samples must be prepared: a set of texts is collected, and each text is assigned a class. These classes are…

Computation and Language · Computer Science 2025-08-26 Aleksandr Tsymbalov , Mikhail Khovrichev

The 2019 WMT Biomedical translation task involved translating Medline abstracts. We approached this using transfer learning to obtain a series of strong neural models on distinct domains, and combining them into multi-domain ensembles. We…

Computation and Language · Computer Science 2019-06-14 Danielle Saunders , Felix Stahlberg , Bill Byrne

We describe an unsupervised method to create pseudo-parallel corpora for machine translation (MT) from unaligned text. We use multilingual BERT to create source and target sentence embeddings for nearest-neighbor search and adapt the model…

Computation and Language · Computer Science 2020-10-16 Phillip Keung , Julian Salazar , Yichao Lu , Noah A. Smith

Prior study has shown that pretrained language models (PLM) can boost the performance of text-based recommendation. In contrast to previous works that either use PLM to encode user history as a whole input text, or impose an additional…

Computation and Language · Computer Science 2023-05-26 Zhiming Mao , Huimin Wang , Yiming Du , Kam-fai Wong

Advances in Large Language Models (e.g., GPT-4, LLaMA) have improved the generation of coherent sentences resembling human writing on a large scale, resulting in the creation of so-called deepfake texts. However, this progress poses…

Computation and Language · Computer Science 2023-10-11 Adaku Uchendu , Jooyoung Lee , Hua Shen , Thai Le , Ting-Hao 'Kenneth' Huang , Dongwon Lee

In this paper, we describe the approach that we employed to address the task of Entity Recognition over Wet Lab Protocols -- a shared task in EMNLP WNUT-2020 Workshop. Our approach is composed of two phases. In the first phase, we…

Computation and Language · Computer Science 2020-12-17 Janvijay Singh , Anshul Wadhawan

This work explores a novel data augmentation method based on Large Language Models (LLMs) for predicting item difficulty and response time of retired USMLE Multiple-Choice Questions (MCQs) in the BEA 2024 Shared Task. Our approach is based…

Computation and Language · Computer Science 2024-04-23 Ana-Cristina Rogoz , Radu Tudor Ionescu

Extracting drug use information from unstructured Electronic Health Records remains a major challenge in clinical Natural Language Processing. While Large Language Models demonstrate advancements, their use in clinical NLP is limited by…

Computation and Language · Computer Science 2026-02-11 Huu-Huy-Hoang Tran , Gia-Bao Duong , Quoc-Viet-Anh Tran , Thi-Hai-Yen Vuong , Hoang-Quynh Le

Separating disinformation from fact on the web has long challenged both the search and the reasoning powers of humans. We show that the reasoning power of large language models (LLMs) and the retrieval power of modern search engines can be…

Computation and Language · Computer Science 2024-11-11 Christopher Malon

Effectively modeling text-rich fresh content such as news articles at document-level is a challenging problem. To ensure a content-based model generalize well to a broad range of applications, it is critical to have a training dataset that…

Computation and Language · Computer Science 2021-06-08 Jialu Liu , Tianqi Liu , Cong Yu

Entity resolution is a widely studied problem with several proposals to match records across relations. Matching textual content is a widespread task in many applications, such as question answering and search. While recent methods achieve…

Databases · Computer Science 2021-12-17 Naser Ahmadi , Hansjorg Sand , Paolo Papotti

Toxicity is pervasive in social media and poses a major threat to the health of online communities. The recent introduction of pre-trained language models, which have achieved state-of-the-art results in many NLP tasks, has transformed the…

Computation and Language · Computer Science 2021-10-11 Erik Yan , Harish Tayyar Madabushi

Sentiment Analysis is a well-studied field of Natural Language Processing. However, the rapid growth of social media and noisy content within them poses significant challenges in addressing this problem with well-established methods and…

Computation and Language · Computer Science 2020-07-28 Soroush Javdan , Taha Shangipour ataei , Behrouz Minaei-Bidgoli

The rapid advancement of Large Language Models (LLMs) has significantly enhanced the capabilities of text generators. With the potential for misuse escalating, the importance of discerning whether texts are human-authored or generated by…

Multimedia · Computer Science 2024-03-12 Travis Munyer , Abdullah Tanvir , Arjon Das , Xin Zhong