English
Related papers

Related papers: A Finnish News Corpus for Named Entity Recognition

200 papers

This paper presents a new annotated corpus of 513 anonymized radiology reports written in Spanish. Reports were manually annotated with entities, negation and uncertainty terms and relations. The corpus was conceived as an evaluation…

Computation and Language · Computer Science 2017-11-01 Viviana Cotik , Darío Filippo , Roland Roller , Hans Uszkoreit , Feiyu Xu

Linking concepts and named entities to knowledge bases has become a crucial Natural Language Understanding task. In this respect, recent works have shown the key advantage of exploiting textual definitions in various Natural Language…

Computation and Language · Computer Science 2017-02-22 José Camacho Collados , Claudio Delli Bovi , Alessandro Raganato , Roberto Navigli

The Novelties corpus is a collection of novels (and parts of novels) annotated for Named Entity Recognition (NER) among other tasks. This document describes the guidelines applied during its annotation. It contains the instructions used by…

Computation and Language · Computer Science 2024-10-07 Arthur Amalvy , Vincent Labatut

We introduce a novel multilingual hierarchical corpus annotated for entity framing and role portrayal in news articles. The dataset uses a unique taxonomy inspired by storytelling elements, comprising 22 fine-grained roles, or archetypes,…

We present the first large scale corpus for entity resolution in email conversations (CEREC). The corpus consists of 6001 email threads from the Enron Email Corpus containing 36,448 email messages and 60,383 entity coreference chains. The…

Computation and Language · Computer Science 2021-06-03 Parag Pravin Dakle , Dan I. Moldovan

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

We introduce KyrgyzNER, the first manually annotated named entity recognition dataset for the Kyrgyz language. Comprising 1,499 news articles from the 24.KG news portal, the dataset contains 10,900 sentences and 39,075 entity mentions…

Computation and Language · Computer Science 2025-09-24 Timur Turatali , Anton Alekseev , Gulira Jumalieva , Gulnara Kabaeva , Sergey Nikolenko

In the current work, we present a description of the system submitted to WMT 2018 News Translation Shared task. The system was created to translate news text from Finnish to English. The system used a Character Based Neural Machine…

Computation and Language · Computer Science 2019-08-02 Sainik Kumar Mahata , Dipankar Das , Sivaji Bandyopadhyay

Nowadays, the rapid diffusion of fake news poses a significant problem, as it can spread misinformation and confusion. This paper aims to develop an advanced machine learning solution for detecting fake news articles. Leveraging a…

Machine Learning · Computer Science 2024-11-19 Tanjina Sultana Camelia , Faizur Rahman Fahim , Md. Musfique Anwar

We present the development of a Named Entity Recognition (NER) dataset for Tagalog. This corpus helps fill the resource gap present in Philippine languages today, where NER resources are scarce. The texts were obtained from a pretraining…

Computation and Language · Computer Science 2023-11-14 Lester James V. Miranda

We present a new open-source parallel corpus consisting of news articles collected from the Bianet magazine, an online newspaper that publishes Turkish news, often along with their translations in English and Kurdish. In this paper, we…

Computation and Language · Computer Science 2018-05-15 Duygu Ataman

There are a lot of tools and resources available for processing Finnish. In this paper, we survey recent papers focusing on Finnish NLP related to many different subcategories of NLP such as parsing, generation, semantics and speech. NLP…

Computation and Language · Computer Science 2021-09-24 Mika Hämäläinen , Khalid Alnajjar

Named Entity Recognition is an information extraction task that serves as a preprocessing step for other natural language processing tasks, such as machine translation, information retrieval, and question answering. Named entity recognition…

Computation and Language · Computer Science 2022-07-05 Ebrahim Chekol Jibril , A. Cüneyd Tantğ

We present a freely available, genre-balanced English web corpus totaling 4M tokens and featuring a large number of high-quality automatic annotation layers, including dependency trees, non-named entity annotations, coreference resolution,…

Computation and Language · Computer Science 2020-06-19 Luke Gessler , Siyao Peng , Yang Liu , Yilun Zhu , Shabnam Behzad , Amir Zeldes

Large-scale news corpora support a wide range of research in Computational Social Science and NLP, yet access remains constrained: commercial archives impose prohibitive costs and licensing restrictions, while open alternatives like Common…

Computation and Language · Computer Science 2026-05-19 Ruggero Marino Lazzaroni , Jana Lasser , Kirill Solovev

In the context of low-resource languages, the Algerian dialect (AD) faces challenges due to the absence of annotated corpora, hindering its effective processing, notably in Machine Learning (ML) applications reliant on corpora for training…

Computation and Language · Computer Science 2024-11-08 Amin Abdedaiem , Abdelhalim Hafedh Dahou , Mohamed Amine Cheragui , Brigitte Mathiak

In this paper, we introduce a dataset of multilingual news articles covering the 2021 Tokyo Olympics. A total of 10,940 news articles were gathered from 1,918 different publishers, covering 1,350 sub-events of the 2021 Olympics, and…

Information Retrieval · Computer Science 2025-02-17 Erik Novak , Erik Calcina , Dunja Mladenić , Marko Grobelnik

Deep learning based natural language processing model is proven powerful, but need large-scale dataset. Due to the significant gap between the real-world tasks and existing Chinese corpus, in this paper, we introduce a large-scale corpus of…

Computation and Language · Computer Science 2018-11-27 Jianyu Zhao , Zhuoran Ji

Named Entity Recognition and Classification (NERC) is a process of identification of proper nouns in the text and classification of those nouns into certain predefined categories like person name, location, organization, date, and time etc.…

Computation and Language · Computer Science 2015-09-21 S. Amarappa , S. V. Sathyanarayana

This Data Descriptor introduces the dataset Enevaeldens Nyheder Online (News during Absolutism Online). The Enevaeldens Nyheder Online (ENO) dataset provides a reconstruction of the contents of major newspapers in Denmark and Norway during…

Digital Libraries · Computer Science 2025-09-03 Johan Heinsen , Camilla Bøgeskov