中文
相关论文

相关论文: A Finnish News Corpus for Named Entity Recognition

200 篇论文

We present a corpus of 5,000 richly annotated abstracts of medical articles describing clinical randomized controlled trials. Annotations include demarcations of text spans that describe the Patient population enrolled, the Interventions…

计算与语言 · 计算机科学 2018-06-13 Benjamin Nye , Junyi Jessy Li , Roma Patel , Yinfei Yang , Iain J. Marshall , Ani Nenkova , Byron C. Wallace

This paper introduces a named entity recognition approach in textual corpus. This Named Entity (NE) can be a named: location, person, organization, date, time, etc., characterized by instances. A NE is found in texts accompanied by…

信息检索 · 计算机科学 2011-03-01 Wahiba Ben Abdessalem Karaa

Morphological analysis (MA) and lexical normalization (LN) are both important tasks for Japanese user-generated text (UGT). To evaluate and compare different MA/LN systems, we have constructed a publicly available Japanese UGT corpus. Our…

计算与语言 · 计算机科学 2021-04-09 Shohei Higashiyama , Masao Utiyama , Taro Watanabe , Eiichiro Sumita

Named Entity Recognition (NER) is an important task in natural language processing that aims to identify and extract key entities from unstructured text. We present a novel application of NER in plasma physics research articles and address…

计算与语言 · 计算机科学 2026-02-13 Muhammad Haris , Hans Höft , Markus M. Becker , Markus Stocker

We present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies. We correct annotation…

计算与语言 · 计算机科学 2025-12-19 Chester Palen-Michel , Maxwell Pickering , Maya Kruse , Jonne Sälevä , Constantine Lignos

This article presents a sentence-level sentiment dataset for the Croatian news domain. In addition to the 3K annotated texts already present, our dataset contains 14.5K annotated sentence occurrences that have been tagged with 5 classes. We…

计算与语言 · 计算机科学 2023-05-16 Gaurish Thakkar , Nives Mikelic Preradović , Marko Tadić

We introduce ParaNames, a massively multilingual parallel name resource consisting of 140 million names spanning over 400 languages. Names are provided for 16.8 million entities, and each entity is mapped from a complex type hierarchy to a…

计算与语言 · 计算机科学 2024-05-16 Jonne Sälevä , Constantine Lignos

This work presents a unified knowledge protocol, called UKnow, which facilitates knowledge-based studies from the perspective of data. Particularly focusing on visual and linguistic modalities, we categorize data knowledge into five unit…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Biao Gong , Shuai Tan , Yutong Feng , Xiaoying Xie , Yuyuan Li , Chaochao Chen , Kecheng Zheng , Yujun Shen , Deli Zhao

Timely analysis of cyber-security information necessitates automated information extraction from unstructured text. While state-of-the-art extraction methods produce extremely accurate results, they require ample training data, which is…

信息检索 · 计算机科学 2014-06-11 Robert A. Bridges , Corinne L. Jones , Michael D. Iannacone , Kelly M. Testa , John R. Goodall

Creating open-domain chatbots requires large amounts of conversational data and related benchmark tasks to evaluate them. Standardized evaluation tasks are crucial for creating automatic evaluation metrics for model development; otherwise,…

计算与语言 · 计算机科学 2020-08-20 Katri Leino , Juho Leinonen , Mittul Singh , Sami Virpioja , Mikko Kurimo

The lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction (GEC). As a complementary new resource for these tasks, we present the GitHub Typo…

计算与语言 · 计算机科学 2019-12-02 Masato Hagiwara , Masato Mita

Most research on emotion analysis from text focuses on the task of emotion classification or emotion intensity regression. Fewer works address emotions as a phenomenon to be tackled with structured learning, which can be explained by the…

计算与语言 · 计算机科学 2020-03-04 Laura Bostan , Evgeny Kim , Roman Klinger

What can we learn from the classics of Finnish literature by using computational emotion analysis? This article tries to answer this question by examining how computational methods of sentiment analysis can be used in the study of literary…

计算与语言 · 计算机科学 2024-06-04 Emily Ohman , Riikka Rossi

High throughput extraction and structured labeling of data from academic articles is critical to enable downstream machine learning applications and secondary analyses. We have embedded multimodal data curation into the academic publishing…

计算与语言 · 计算机科学 2024-09-26 Jorge Abreu-Vicente , Hannah Sonntag , Thomas Eidens , Cassie S. Mitchell , Thomas Lemberger

We present the Knesset Corpus, a corpus of Hebrew parliamentary proceedings containing over 30 million sentences (over 384 million tokens) from all the (plenary and committee) protocols held in the Israeli parliament between 1998 and 2022.…

计算与语言 · 计算机科学 2025-06-02 Gili Goldin , Nick Howell , Noam Ordan , Ella Rabinovich , Shuly Wintner

News articles, image captions, product reviews and many other texts mention people and organizations whose name recognition could vary for different audiences. In such cases, background information about the named entities could be provided…

计算与语言 · 计算机科学 2020-11-09 Yova Kementchedjhieva , Di Lu , Joel Tetreault

The most common Named Entity Recognizers are usually sequence taggers trained on fully annotated corpora, i.e. the class of all words for all entities is known. Partially annotated corpora, i.e. some but not all entities of some types are…

计算与语言 · 计算机科学 2022-04-21 Michael Strobl , Amine Trabelsi , Osmar Zaiane

News article revision histories have the potential to give us novel insights across varied fields of linguistics and social sciences. In this work, we present, to our knowledge, the first publicly available dataset of news article revision…

计算与语言 · 计算机科学 2022-07-01 Alexander Spangher , Jonathan May

Having a quality annotated corpus is essential especially for applied research. Despite the recent focus of Web science community on researching about cyberbullying, the community dose not still have standard benchmarks. In this paper, we…

This paper presents Wojood, a corpus for Arabic nested Named Entity Recognition (NER). Nested entities occur when one entity mention is embedded inside another entity mention. Wojood consists of about 550K Modern Standard Arabic (MSA) and…

计算与语言 · 计算机科学 2022-05-24 Mustafa Jarrar , Mohammed Khalilia , Sana Ghanem