English
Related papers

Related papers: Guidelines for the Creation of an Annotated Corpus

200 papers

In this paper we present a corpus of Russian strategic planning documents, RuREBus. This project is grounded both from language technology and e-government perspectives. Not only new language sources and tools are being developed, but also…

Automatic medical text simplification can assist providers with patient-friendly communication and make medical texts more accessible, thereby improving health literacy. But curating a quality corpus for this task requires the supervision…

Computation and Language · Computer Science 2023-02-21 Chandrayee Basu , Rosni Vasu , Michihiro Yasunaga , Qian Yang

Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, traditionally obtained by manually designing an annotation schema and exhaustively labeling the…

Computation and Language · Computer Science 2026-04-13 Shahar Levy , Eliya Habba , Reshef Mintz , Barak Raveh , Renana Keydar , Gabriel Stanovsky

Multi-document summarization (MDS) is the task of reflecting key points from any set of documents into a concise text paragraph. In the past, it has been used to aggregate news, tweets, product reviews, etc. from various sources. Owing to…

Computation and Language · Computer Science 2020-10-06 Alvin Dey , Tanya Chowdhury , Yash Kumar Atri , Tanmoy Chakraborty

This document provides extensive guidelines and examples for Rhetorical Structure Theory (RST) annotation in Mandarin Chinese. The guideline is divided into three sections. We first introduce preprocessing steps to prepare data for RST…

Computation and Language · Computer Science 2022-12-13 Siyao Peng , Yang Janet Liu , Amir Zeldes

Social Media platforms have offered invaluable opportunities for linguistic research. The availability of up-to-date data, coming from any part in the world, and coming from natural contexts, has allowed researchers to study language in…

Computation and Language · Computer Science 2024-07-23 Simon Gonzalez

The evolution of AI systems toward agentic operation and context-aware retrieval necessitates transforming unstructured text into structured formats like tables, knowledge graphs, and charts. While such conversions enable critical…

Computation and Language · Computer Science 2025-08-19 Zheye Deng , Chunkit Chan , Tianshi Zheng , Wei Fan , Weiqi Wang , Yangqiu Song

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

The OpenCitations organization is working on ingesting citation data and bibliographic metadata directly provided by the community (e.g., scholars and publishers). The aim is to improve the general coverage of open citations, which is still…

Digital Libraries · Computer Science 2022-09-26 Arcangelo Massari , Ivan Heibi

Most research on emotion analysis from text focuses on the task of emotion classification or emotion intensity regression. Fewer works address emotions as a phenomenon to be tackled with structured learning, which can be explained by the…

Computation and Language · Computer Science 2020-03-04 Laura Bostan , Evgeny Kim , Roman Klinger

The paper introduces a framework for representation and acquisition of knowledge emerging from large samples of textual data. We utilise a tensor-based, distributional representation of simple statements extracted from text, and show how…

Artificial Intelligence · Computer Science 2012-10-12 Vit Novacek

This paper presents how the online tool GREW-MATCH can be used to make queries and visualise data from existing semantically annotated corpora. A dedicated syntax is available to construct simple to complex queries and execute them against…

Artificial Intelligence · Computer Science 2022-07-26 Maxime Amblard , Bruno Guillaume , Siyana Pavlova , Guy Perrier

Large-scale datasets are essential to modern day deep learning. Advocates argue that understanding these methods requires dataset transparency (e.g. "dataset curation, motivation, composition, collection process, etc..."). However, almost…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Nadine Chang , Francesco Ferroni , Michael J. Tarr , Martial Hebert , Deva Ramanan

This paper presents a scheme for annotating coreference across news articles, extending beyond traditional identity relations by also considering near-identity and bridging relations. It includes a precise description of how to set up…

Computation and Language · Computer Science 2023-10-19 Jakob Vogel

Annotations play a vital role in highlighting critical aspects of visualizations, aiding in data externalization and exploration, collaborative sensemaking, and visual storytelling. However, despite their widespread use, we identified a…

Human-Computer Interaction · Computer Science 2026-04-10 Md Dilshadur Rahman , Ghulam Jilani Quadri , Bhavana Doppalapudi , Danielle Albers Szafir , Paul Rosen

In this paper, we present a semantic web approach for modelling the process of creating new technical and regulatory documents related to the Building sector. This industry, among other industries, is currently experiencing a phenomenal…

Information Retrieval · Computer Science 2013-02-20 Khalil Riad Bouzidi , Bruno Fies , Marc Bourdeau , Catherine Faron-Zucker , Nhan Le-Thanh

This technical report presents an evaluation of the ontology annotations in the metadata of a subset of entries of MetaboLights, a database for Metabolomics experiments and derived information. The work includes a manual analysis of the…

Digital Libraries · Computer Science 2016-04-29 Camila Ramos , Marco Louro , Miguel Santos , Francisco M. Couto

Named entity recognition (NER) is the very first step in the linguistic processing of any new domain. It is currently a common process in BioNLP on English clinical text. However, it is still in its infancy in other major languages, as it…

Computation and Language · Computer Science 2019-12-20 Fernando Sánchez León , Ana González Ledesma

A novel approach to the fully automated, unsupervised extraction of dependency grammars and associated syntax-to-semantic-relationship mappings from large text corpora is described. The suggested approach builds on the authors' prior work…

Computation and Language · Computer Science 2014-01-16 Linas Vepstas , Ben Goertzel

The Annotation Graph Toolkit (AGTK) is a collection of software which facilitates development of linguistic annotation tools. AGTK provides a database interface which allows applications to use a database server for persistent storage. This…

Computation and Language · Computer Science 2007-05-23 Xiaoyi Ma , Haejoong Lee , Steven Bird , Kazuaki Maeda
‹ Prev 1 8 9 10 Next ›