English
Related papers

Related papers: A Finnish News Corpus for Named Entity Recognition

200 papers

There is an increasing interest in studying natural language and computer code together, as large corpora of programming texts become readily available on the Internet. For example, StackOverflow currently has over 15 million programming…

Computation and Language · Computer Science 2020-11-17 Jeniya Tabassum , Mounica Maddela , Wei Xu , Alan Ritter

This article presents the application of the Universal Named Entity framework to generate automatically annotated corpora. By using a workflow that extracts Wikipedia data and meta-data and DBpedia information, we generated an English…

Computation and Language · Computer Science 2022-12-15 Diego Alves , Gaurish Thakkar , Marko Tadić

Identification of named entities from legal texts is an essential building block for developing other legal Artificial Intelligence applications. Named Entities in legal texts are slightly different and more fine-grained than commonly used…

Computation and Language · Computer Science 2022-11-08 Prathamesh Kalamkar , Astha Agarwal , Aman Tiwari , Smita Gupta , Saurabh Karn , Vivek Raghavan

Fake news detection is a challenging task aiming to reduce human time and effort to check the truthfulness of news. Automated approaches to combat fake news, however, are limited by the lack of labeled benchmark datasets, especially in…

Computation and Language · Computer Science 2021-03-02 Inna Vogel , Jeong-Eun Choi , Meghana Meghana

The scarcity of available text corpora for low-resource languages like Albanian is a serious hurdle for research in natural language processing tasks. This paper introduces AlbNews, a collection of 600 topically labeled news headlines and…

Computation and Language · Computer Science 2024-02-07 Erion Çano , Dario Lamaj

Named entity recognition identifies common classes of entities in text, but these entity labels are generally sparse, limiting utility to downstream tasks. In this work we present ScienceExamCER, a densely-labeled semantic classification…

Computation and Language · Computer Science 2019-11-26 Hannah Smith , Zeyu Zhang , John Culnan , Peter Jansen

We describe a gold standard corpus of protest events that comprise of various local and international sources from various countries in English. The corpus contains document, sentence, and token level annotations. This corpus facilitates…

Computation and Language · Computer Science 2020-08-04 Ali Hürriyetoğlu , Erdem Yörük , Deniz Yüret , Osman Mutlu , Çağrı Yoltar , Fırat Duruşan , Burak Gürel

This paper presents a corpus manually annotated with named entities for six Slavic languages - Bulgarian, Czech, Polish, Slovenian, Russian, and Ukrainian. This work is the result of a series of shared tasks, conducted in 2017-2023 as a…

Computation and Language · Computer Science 2024-04-09 Jakub Piskorski , Michał Marcińczuk , Roman Yangarber

Recognizing non-standard entity types and relations, such as B2B products, product classes and their producers, in news and forum texts is important in application areas such as supply chain monitoring and market research. However, there is…

Computation and Language · Computer Science 2020-04-08 Saskia Schön , Veselina Mironova , Aleksandra Gabryszak , Leonhard Hennig

We present a new corpus comprising annotations of medical entities in case reports, originating from PubMed Central's open access library. In the case reports, we annotate cases, conditions, findings, factors and negation modifiers.…

Computation and Language · Computer Science 2020-03-31 Sarah Schulz , Jurica Ševa , Samuel Rodriguez , Malte Ostendorff , Georg Rehm

News agencies publish news on their websites all over the world. Moreover, creating novel corpuses is necessary to bring natural processing to new domains. Textual processing of online news is challenging in terms of the strategy of…

Computation and Language · Computer Science 2018-08-22 Mohammad Kamel , Hadi Sadoghi-Yazdi

We present the Verifee Dataset: a novel dataset of news articles with fine-grained trustworthiness annotations. We develop a detailed methodology that assesses the texts based on their parameters encompassing editorial transparency,…

Computation and Language · Computer Science 2022-12-19 Matyáš Boháček , Michal Bravanský , Filip Trhlík , Václav Moravec

Although pre-trained named entity recognition (NER) models are highly accurate on modern corpora, they underperform on historical texts due to differences in language OCR errors. In this work, we develop a new NER corpus of 3.6M sentences…

Computation and Language · Computer Science 2023-06-08 Vít Novotný , Kristýna Luger , Michal Štefánik , Tereza Vrabcová , Aleš Horák

Named entities in text documents are the names of people, organization, location or other types of objects in the documents that exist in the real world. A persisting research challenge is to use computational techniques to identify such…

Computation and Language · Computer Science 2019-07-09 Abdulkareem Alsudais , Hovig Tchalian

We present Namesakes, a dataset of ambiguously named entities obtained from English-language Wikipedia and news articles. It consists of 58862 mentions of 4148 unique entities and their namesakes: 1000 mentions from news, 28843 from…

Computation and Language · Computer Science 2021-11-23 Oleg Vasilyev , Aysu Altun , Nidhi Vyas , Vedant Dharnidharka , Erika Lam , John Bohannon

News articles such as sports game reports are often thought to closely follow the underlying game statistics, but in practice they contain a notable amount of background knowledge, interpretation, insight into the game, and quotes that are…

Computation and Language · Computer Science 2019-10-07 Jenna Kanerva , Samuel Rönnqvist , Riina Kekki , Tapio Salakoski , Filip Ginter

Information resources such as newspapers have produced unstructured text data in various languages related to the corona outbreak since December 2019. Analyzing these unstructured texts is time-consuming without representing them in a…

Computation and Language · Computer Science 2024-04-25 Sefika Efeoglu , Adrian Paschke

Entities like person, location, organization are important for literary text analysis. The lack of annotated data hinders the progress of named entity recognition (NER) in literary domain. To promote the research of literary NER, we build…

Computation and Language · Computer Science 2024-10-16 Hanjie Zhao , Jinge Xie , Yuchen Yan , Yuxiang Jia , Yawen Ye , Hongying Zan

This paper introduces the Multi-Genre Natural Language Inference (MultiNLI) corpus, a dataset designed for use in the development and evaluation of machine learning models for sentence understanding. In addition to being one of the largest…

Computation and Language · Computer Science 2018-02-21 Adina Williams , Nikita Nangia , Samuel R. Bowman

Identifying named entities such as a person, location or organization, in documents can highlight key information to readers. Training Named Entity Recognition (NER) models requires an annotated data set, which can be a time-consuming…

Computation and Language · Computer Science 2022-12-20 Ting Wai Terence Au , Ingemar J. Cox , Vasileios Lampos