English
Related papers

Related papers: Guidelines for the Creation of an Annotated Corpus

200 papers

We describe a formal model for annotating linguistic artifacts, from which we derive an application programming interface (API) to a suite of tools for manipulating these annotations. The abstract logical model provides for a range of…

Computation and Language · Computer Science 2007-05-23 Steven Bird , David Day , John Garofolo , John Henderson , Christophe Laprun , Mark Liberman

This work distinguishes between translated and original text in the UN protocol corpus. By modeling the problem as classification problem, we can achieve up to 95% classification accuracy. We begin by deriving a parallel corpus for…

Computation and Language · Computer Science 2018-05-22 Elad Tolochinsky , Ohad Mosafi , Ella Rabinovich , Shuly Wintner

The article examines the theoretical, methodological, and technical foundations of research on audiovisual corpora within the field of digital humanities. It outlines the main transversal issues underlying the processes of constructing,…

Digital Libraries · Computer Science 2025-11-07 Peter Stockinger

Research in Computational Linguistics is dependent on text corpora for training and testing new tools and methodologies. While there exists a plethora of annotated linguistic information, these corpora are often not interoperable without…

Computation and Language · Computer Science 2020-11-03 Timo Lek , Anna de Groot , Tobias Kuhn , Roser Morante

Literary texts are usually rich in meanings and their interpretation complicates corpus studies and automatic processing. There have been several attempts to create collections of literary texts with annotation of literary elements like the…

Computation and Language · Computer Science 2021-11-16 Elena Mikhalkova , Timofei Protasov , Anastasiia Drozdova , Anastasiia Bashmakova , Polina Gavin

In terms of annotation structure, most learner corpora rely on holistic flat label inventories which, even when extensive, do not explicitly separate multiple linguistic dimensions. This makes linguistically deep annotation difficult and…

Computation and Language · Computer Science 2026-02-04 Elif Sayar , Tolgahan Türker , Anna Golynskaia Knezhevich , Bihter Dereli , Ayşe Demirhas , Lionel Nicolas , Gülşen Eryiğit

We introduce the Constructive Comments Corpus (C3), comprised of 12,000 annotated news comments, intended to help build new tools for online communities to improve the quality of their discussions. We define constructive comments as…

Computation and Language · Computer Science 2020-08-06 Varada Kolhatkar , Nithum Thain , Jeffrey Sorensen , Lucas Dixon , Maite Taboada

Current approaches to the annotation process focus on annotation schemas, languages for annotation, or are very application driven. In this paper it is proposed that a more flexible architecture for annotation requires a knowledge component…

Digital Libraries · Computer Science 2007-05-23 Afzal Ballim , Nastaran Fatemi , Hatem Ghorbel , Vincenzo Pallotta

Learned vector representations of words are useful tools for many information retrieval and natural language processing tasks due to their ability to capture lexical semantics. However, while many such tasks involve or even rely on named…

Computation and Language · Computer Science 2020-02-13 Satya Almasian , Andreas Spitz , Michael Gertz

We present a method for learning large-scale, broad-coverage construction grammars from corpora of language use. Starting from utterances annotated with constituency structure and semantic frames, the method facilitates the learning of…

Computation and Language · Computer Science 2026-05-27 Paul Van Eecke , Katrien Beuls

This paper elaborates on the notion of uncertainty in the context of annotation in large text corpora, specifically focusing on (but not limited to) historical languages. Such uncertainty might be due to inherent properties of the language,…

Computation and Language · Computer Science 2021-05-31 Marie-Luis Merten , Marcel Wever , Michaela Geierhos , Doris Tophinke , Eyke Hüllermeier

This paper presents a new challenging information extraction task in the domain of materials science. We develop an annotation scheme for marking information on experiments related to solid oxide fuel cells in scientific publications, such…

Computation and Language · Computer Science 2020-06-05 Annemarie Friedrich , Heike Adel , Federico Tomazic , Johannes Hingerl , Renou Benteau , Anika Maruscyk , Lukas Lange

While annotated corpora are crucial in the field of natural language processing (NLP), those containing copyrighted material are difficult to exchange among researchers. Yet, such corpora are necessary to fully represent the diversity of…

Computation and Language · Computer Science 2026-04-28 Arthur Amalvy , Vincent Labatut , Xavier Bost , Hen-Hsen Huang

We describe Artemis (Annotation methodology for Rich, Tractable, Extractive, Multi-domain, Indicative Summarization), a novel hierarchical annotation process that produces indicative summaries for documents from multiple domains. Current…

Computation and Language · Computer Science 2020-05-15 Rahul Jha , Keping Bi , Yang Li , Mahdi Pakdaman , Asli Celikyilmaz , Ivan Zhiboedov , Kieran McDonald

Many communities, including the scientific community, develop implicit writing norms. Understanding them is crucial for effective communication with that community. Writers gradually develop an implicit understanding of norms by reading…

Human-Computer Interaction · Computer Science 2025-03-18 Hai Dang , Chelse Swoopes , Daniel Buschek , Elena L. Glassman

Large-scale, high-quality corpora are critical for advancing research in coreference resolution. However, existing datasets vary in their definition of coreferences and have been collected via complex and lengthy guidelines that are curated…

Computation and Language · Computer Science 2022-10-14 Ankita Gupta , Marzena Karpinska , Wenlong Zhao , Kalpesh Krishna , Jack Merullo , Luke Yeh , Mohit Iyyer , Brendan O'Connor

Academic research is an exploratory activity to discover new solutions to problems. By this nature, academic research works perform literature reviews to distinguish their novelties from prior work. In natural language processing, this…

Computation and Language · Computer Science 2025-05-19 Xiangci Li , Biswadip Mandal , Jessica Ouyang

This paper presents Charon, a web tool for annotating multimodal corpora with FrameNet categories. Annotation can be made for corpora containing both static images and video sequences paired - or not - with text sequences. The pipeline…

Computation and Language · Computer Science 2022-05-25 Frederico Belcavello , Marcelo Viridiano , Ely Edison Matos , Tiago Timponi Torrent

Unstructured information comprises a valuable source of data in clinical records. For text mining in clinical records, concept extraction is the first step in finding assertions and relationships. This study presents a system developed for…

Information Retrieval · Computer Science 2010-12-09 Ning Kang , Rogier Barendse , Zubair Afzal , Bharat Singh , Martijn J. Schuemie , Erik M. van Mulligen , Jan A. Kors

Abstract. When writing an academic paper, researchers often spend considerable time reviewing and summarizing papers to extract relevant citations and data to compose the Introduction and Related Work sections. To address this problem, we…

Information Retrieval · Computer Science 2023-06-22 Juan Ramirez-Orta , Eduardo Xamena , Ana Maguitman , Axel J. Soto , Flavia P. Zanoto , Evangelos Milios