English
Related papers

Related papers: RuCoCo: a new Russian corpus with coreference anno…

200 papers

We describe a recently developed corpus annotation scheme for evaluating parsers that avoids shortcomings of current methods. The scheme encodes grammatical relations between heads and dependents, and has been used to mark up a new…

Computation and Language · Computer Science 2007-05-23 John Carroll , Guido Minnen , Ted Briscoe

We present a comprehensive corpus of Russian primary and secondary legislation adopted between 1991 and 2025, comprising 304,382 texts (194,425,905 tokens). The corpus is available in two versions: the basic version contains texts with…

Computation and Language · Computer Science 2026-04-29 Denis Saveliev , Ruslan Kuchakov

The development of large and super-large language models, such as GPT-3, T5, Switch Transformer, ERNIE, etc., has significantly improved the performance of text generation. One of the important research directions in this area is the…

Computation and Language · Computer Science 2022-06-22 Valeriya Goloviznina , Evgeny Kotelnikov

Corpora that contain tabular data such as WebTables are a vital resource for the academic community. Essentially, they are the backbone of any modern research in information management. They are used for various tasks of data extraction,…

Computation and Language · Computer Science 2022-10-13 Platon Fedorov , Alexey Mironov , George Chernishev

The Russian Drug Reaction Corpus (RuDReC) is a new partially annotated corpus of consumer reviews in Russian about pharmaceutical products for the detection of health-related named entities and the effectiveness of pharmaceutical products.…

Computation and Language · Computer Science 2023-11-21 Elena Tutubalina , Ilseyar Alimova , Zulfat Miftahutdinov , Andrey Sakhovskiy , Valentin Malykh , Sergey Nikolenko

We present a freely available, genre-balanced English web corpus totaling 4M tokens and featuring a large number of high-quality automatic annotation layers, including dependency trees, non-named entity annotations, coreference resolution,…

Computation and Language · Computer Science 2020-06-19 Luke Gessler , Siyao Peng , Yang Liu , Yilun Zhu , Shabnam Behzad , Amir Zeldes

This paper presents a corpus annotated for the task of direct-speech extraction in Croatian. The paper focuses on the annotation of the quotation, co-reference resolution, and sentiment annotation in SETimes news corpus in Croatian and on…

Computation and Language · Computer Science 2022-12-15 Jelena Sarajlić , Gaurish Thakkar , Diego Alves , Nives Mikelic Preradović

Large sense-annotated datasets are increasingly necessary for training deep supervised systems in Word Sense Disambiguation. However, gathering high-quality sense-annotated data for as many instances as possible is a laborious and expensive…

Computation and Language · Computer Science 2020-03-16 Tommaso Pasini , Jose Camacho-Collados

We describe a gold standard corpus of protest events that comprise of various local and international sources from various countries in English. The corpus contains document, sentence, and token level annotations. This corpus facilitates…

Computation and Language · Computer Science 2020-08-04 Ali Hürriyetoğlu , Erdem Yörük , Deniz Yüret , Osman Mutlu , Çağrı Yoltar , Fırat Duruşan , Burak Gürel

The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered from issues such as opaque process, difficulty of…

Computation and Language · Computer Science 2025-09-22 Qiuyang Lu , Fangjian Shen , Zhengkai Tang , Qiang Liu , Hexuan Cheng , Hui Liu , Wushao Wen

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G…

Computation and Language · Computer Science 2020-03-06 Liang Xu , Xuanwei Zhang , Qianqian Dong

In this paper, we present a distributional word embedding model trained on one of the largest available Russian corpora: Araneum Russicum Maximum (over 10 billion words crawled from the web). We compare this model to the model trained on…

Computation and Language · Computer Science 2018-01-22 Andrey Kutuzov , Maria Kunilovskaya

Scientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article's impacts on research community. This paper provides novel solutions to these two challenges. We…

Computation and Language · Computer Science 2019-09-17 Michihiro Yasunaga , Jungo Kasai , Rui Zhang , Alexander R. Fabbri , Irene Li , Dan Friedman , Dragomir R. Radev

We describe a new sense-tagged corpus for word sense disambiguation. The corpus is constituted of instances of 20 French polysemous verbs. Each verb instance is annotated with three sense labels: (1) the actual translation of the verb in…

Computation and Language · Computer Science 2026-05-28 Myriam Rakho , Eric Laporte , Matthieu Constant

Developing natural language processing tools for clinical text requires annotated datasets, yet French oncology resources remain scarce. We present FRACCO (FRench Annotated Corpus for Clinical Oncology) an expert-annotated corpus of 1301…

Analyzing how humans revise their writings is an interesting research question, not only from an educational perspective but also in terms of artificial intelligence. Better understanding of this process could facilitate many NLP…

Computation and Language · Computer Science 2022-06-06 Omid Kashefi , Tazin Afrin , Meghan Dale , Christopher Olshefski , Amanda Godley , Diane Litman , Rebecca Hwa

This paper is devoted to the study of methods for information extraction (entity recognition and relation classification) from scientific texts on information technology. Scientific publications provide valuable information into…

Computation and Language · Computer Science 2020-12-29 Elena Bruches , Alexey Pauls , Tatiana Batura , Vladimir Isachenko

Automatic summarization techniques aim to shorten and generalize information given in the text while preserving its core message and the most relevant ideas. This task can be approached and treated with a variety of methods, however, not…

Computation and Language · Computer Science 2021-08-10 Alexandr Nikolich , Arina Puchkova

This paper provides a comprehensive overview of the gapping dataset for Russian that consists of 7.5k sentences with gapping (as well as 15k relevant negative sentences) and comprises data from various genres: news, fiction, social media…

Computation and Language · Computer Science 2019-06-11 Maria Ponomareva , Kira Droganova , Ivan Smurov , Tatiana Shavrina

In this paper, we introduce the Dialogue Evaluation shared task on extraction of structured opinions from Russian news texts. The task of the contest is to extract opinion tuples for a given sentence; the tuples are composed of a sentiment…

Computation and Language · Computer Science 2025-04-10 Natalia Loukachevitch , Natalia Tkachenko , Anna Lapanitsyna , Mikhail Tikhomirov , Nicolay Rusnachenko