English
Related papers

Related papers: Guidelines and a Corpus for Extracting Biographica…

200 papers

Quotation extraction is a widely useful task both from a sociological and from a Natural Language Processing perspective. However, very little data is available to study this task in languages other than English. In this paper, we present a…

Computation and Language · Computer Science 2023-09-20 Ange Richard , Laura Alonzo-Canul , François Portet

We report on a language resource consisting of 2000 annotated bibliography entries, which is being analyzed as part of our research on indicative document summarization. We show how annotated bibliographies cover certain aspects of…

Computation and Language · Computer Science 2007-05-23 Min-Yen Kan , Judith L. Klavans , Kathleen R. McKeown

The schema.org initiative led by the four major search engines curates a vocabulary for describing web content. The number of semantic annotations on the web are increasing, mostly due to the industrial incentives provided by those search…

Information Retrieval · Computer Science 2018-05-16 Umutcan Şimşek , Elias Kärle , Dieter Fensel

Thanks to the rise of wearable and connected devices, sensor-generated time series comprise a large and growing fraction of the world's data. Unfortunately, extracting value from this data can be challenging, since sensors report low-level…

Machine Learning · Statistics 2016-09-30 Davis W. Blalock , John V. Guttag

In this paper, we present CrudeOilNews, a corpus of English Crude Oil news for event extraction. It is the first of its kind for Commodity News and serve to contribute towards resource building for economic and financial text mining. This…

Computation and Language · Computer Science 2022-04-11 Meisin Lee , Lay-Ki Soon , Eu-Gene Siew , Ly Fie Sugianto

When speaking or writing, people omit information that seems clear and evident, such that only part of the message is expressed in words. Especially in argumentative texts it is very common that (important) parts of the argument are implied…

Computation and Language · Computer Science 2019-12-24 Maria Becker , Katharina Korfhage , Anette Frank

Event extraction identifies the central aspects of events from text. It supports event understanding and analysis, which is crucial for tasks such as informed decision-making in emergencies. Therefore, it is necessary to develop automated…

Computation and Language · Computer Science 2026-04-24 Praval Sharma , Ashok Samal , Leen-Kiat Soh , Deepti Joshi

Manually annotated data is key to developing text-mining and information-extraction algorithms. However, human annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to…

Human-Computer Interaction · Computer Science 2020-04-27 Rezarta Islamaj , Dongseop Kwon , Sun Kim , Zhiyong Lu

Event Argument extraction refers to the task of extracting structured information from unstructured text for a particular event of interest. The existing works exhibit poor capabilities to extract causal event arguments like Reason and…

Computation and Language · Computer Science 2021-05-04 Debanjana Kar , Sudeshna Sarkar , Pawan Goyal

Annotating text data for event information extraction systems is hard, expensive, and error-prone. We investigate the feasibility of integrating coarse-grained data (document or sentence labels), which is far more feasible to obtain,…

Computation and Language · Computer Science 2022-05-12 Osman Mutlu

Biomedical researchers use ontologies to annotate their data with ontology terms, enabling better data integration and interoperability. However, the number, variety and complexity of current biomedical ontologies make it cumbersome for…

Artificial Intelligence · Computer Science 2017-06-09 Marcos Martinez-Romero , Clement Jonquet , Martin J. O'Connor , John Graybeal , Alejandro Pazos , Mark A. Musen

Annotating temporal relations (TempRel) between events described in natural language is known to be labor intensive, partly because the total number of TempRels is quadratic in the number of events. As a result, only a small number of…

Computation and Language · Computer Science 2018-04-26 Qiang Ning , Zhongzhi Yu , Chuchu Fan , Dan Roth

Many recent works aim at developing methods and tools for the processing of semantic Web services. In order to be properly tested, these tools must be applied to an appropriate benchmark, taking the form of a collection of semantic WS…

Software Engineering · Computer Science 2015-02-04 Cihan Aksoy , Vincent Labatut , Chantal Cherifi , Jean-François Santucci

Having a quality annotated corpus is essential especially for applied research. Despite the recent focus of Web science community on researching about cyberbullying, the community dose not still have standard benchmarks. In this paper, we…

Computation and Language · Computer Science 2018-05-25 Mohammadreza Rezvan , Saeedeh Shekarpour , Lakshika Balasuriya , Krishnaprasad Thirunarayan , Valerie Shalin , Amit Sheth

Large language models (LLMs) have shown promise for automated text annotation, raising hopes that they might accelerate cross-cultural research by extracting structured data from ethnographic texts. We evaluated 7 state-of-the-art LLMs on…

Computation and Language · Computer Science 2026-01-21 Leonardo S. Goodall , Dor Shilton , Daniel A. Mullins , Harvey Whitehouse

Social biases on Wikipedia, a widely-read global platform, could greatly influence public opinion. While prior research has examined man/woman gender bias in biography articles, possible influences of other demographic attributes limit…

Computation and Language · Computer Science 2022-02-10 Anjalie Field , Chan Young Park , Kevin Z. Lin , Yulia Tsvetkov

We study continual event extraction, which aims to extract incessantly emerging event information while avoiding forgetting. We observe that the semantic confusion on event types stems from the annotations of the same text being updated…

Computation and Language · Computer Science 2023-10-25 Zitao Wang , Xinyi Wang , Wei Hu

In this paper, we present a manually annotated corpus of 10,000 tweets containing public reports of five COVID-19 events, including positive and negative tests, deaths, denied access to testing, claimed cures and preventions. We designed…

Computation and Language · Computer Science 2022-09-12 Shi Zong , Ashutosh Baheti , Wei Xu , Alan Ritter

We present a system that allows life-science researchers to search a linguistically annotated corpus of scientific texts using patterns over dependency graphs, as well as using patterns over token sequences and a powerful variant of boolean…

Computation and Language · Computer Science 2020-06-09 Hillel Taub-Tabib , Micah Shlain , Shoval Sadde , Dan Lahav , Matan Eyal , Yaara Cohen , Yoav Goldberg

Microblogs such as Twitter represent a powerful source of information. Part of this information can be aggregated beyond the level of individual posts. Some of this aggregated information is referring to events that could or should be acted…

Computation and Language · Computer Science 2020-08-04 Ali Hürriyetoğlu