中文
相关论文

相关论文: A Semi-Automated Approach for Information Extracti…

200 篇论文

With rise of digital age, there is an explosion of information in the form of news, articles, social media, and so on. Much of this data lies in unstructured form and manually managing and effectively making use of it is tedious, boring and…

计算与语言 · 计算机科学 2018-07-09 Sonit Singh

In this paper, we present our approach to extracting structured information from unstructured Electronic Health Records (EHR) [2] which can be used to, for example, study adverse drug reactions in patients due to chemicals in their…

计算与语言 · 计算机科学 2020-01-30 Amogh Kamat Tarcar , Aashis Tiwari , Vineet Naique Dhaimodker , Penjo Rebelo , Rahul Desai , Dattaraj Rao

Availability, collection and access to quantitative data, as well as its limitations, often make qualitative data the resource upon which development programs heavily rely. Both traditional interview data and social media analysis can…

计算与语言 · 计算机科学 2017-09-19 Philipp Broniecki , Anna Hanchar , Slava J. Mikhaylov

Data citations provide a foundation for studying research data impact. Collecting and managing data citations is a new frontier in archival science and scholarly communication. However, the discovery and curation of research data citations…

数字图书馆 · 计算机科学 2022-03-11 Lizhou Fan , Sara Lafia , David Bleckley , Elizabeth Moss , Andrea Thomer , Libby Hemphill

This is a machine learning application paper involving big data. We present high-accuracy prediction methods of rare events in semi-structured machine log files, which are produced at high velocity and high volume by NORC's…

机器学习 · 计算机科学 2015-10-06 Sou-Cheng T. Choi

Language is the medium for many political activities, from campaigns to news reports. Natural language processing (NLP) uses computational tools to parse text into key information that is needed for policymaking. In this chapter, we…

计算与语言 · 计算机科学 2023-02-08 Zhijing Jin , Rada Mihalcea

We address the problem of extracting structured representations of economic events from a large corpus of news articles, using a combination of natural language processing and machine learning techniques. The developed techniques allow for…

信息检索 · 计算机科学 2017-09-19 Jan R. Benetka , Krisztian Balog , Kjetil Nørvåg

This paper introduces ENEIDE (Extracting Named Entities from Italian Digital Editions), a silver standard dataset for Named Entity Recognition and Linking (NERL) in historical Italian texts. The corpus comprises 2,111 documents with over…

计算与语言 · 计算机科学 2026-04-01 Cristian Santini , Sebastian Barzaghi , Paolo Sernani , Emanuele Frontoni , Laura Melosi , Mehwish Alam

Our research investigates how Natural Language Processing (NLP) can be used to extract main topics from a larger corpus of written data, as applied to the case of identifying signaling themes in Presidential Directives (PDs) from the Reagan…

计算与语言 · 计算机科学 2025-11-14 C. LeMay , A. Lane , J. Seales , M. Winstead , S. Baty

The progressive digitization of historical archives provides new, often domain specific, textual resources that report on facts and events which have happened in the past; among these, memoirs are a very common type of primary source. In…

计算与语言 · 计算机科学 2021-02-25 Marco Rovera , Federico Nanni , Simone Paolo Ponzetto

Natural language processing tools have become frequently used in social sciences such as economics, political science, and sociology. Many publications apply topic modeling to elicit latent topics in text corpora and their development over…

综合经济学 · 经济学 2024-04-30 W. Benedikt Schmal

We present an extension of sparse PCA, or sparse dictionary learning, where the sparsity patterns of all dictionary elements are structured and constrained to belong to a prespecified set of shapes. This \emph{structured sparse PCA} is…

机器学习 · 统计学 2009-09-09 Rodolphe Jenatton , Guillaume Obozinski , Francis Bach

The package cleanNLP provides a set of fast tools for converting a textual corpus into a set of normalized tables. The underlying natural language processing pipeline utilizes Stanford's CoreNLP library, exposing a number of annotation…

计算与语言 · 计算机科学 2018-05-04 Taylor Arnold

Extracting biographical information from online documents is a popular research topic among the information extraction (IE) community. Various natural language processing (NLP) techniques such as text classification, text summarisation and…

信息检索 · 计算机科学 2022-05-03 Alistair Plum , Tharindu Ranasinghe , Spencer Jones , Constantin Orasan , Ruslan Mitkov

Current social science efforts automatically populate event databases of "who did what to whom?" tuples, by applying event extraction (EE) to text such as news. The event databases are used to analyze sociopolitical dynamics between actor…

计算与语言 · 计算机科学 2024-06-04 Erica Cai , Brendan O'Connor

In this paper we show how to process the NOTAM (Notice to Airmen) data of the field in civil aviation. The main research contents are as follows: 1.Data preprocessing: For the original data of the NOTAM, there is a mixture of Chinese and…

计算与语言 · 计算机科学 2021-06-15 YiPeng Deng , YinHui Luo

Annotating temporal relations (TempRel) between events described in natural language is known to be labor intensive, partly because the total number of TempRels is quadratic in the number of events. As a result, only a small number of…

计算与语言 · 计算机科学 2018-04-26 Qiang Ning , Zhongzhi Yu , Chuchu Fan , Dan Roth

A named entity recognition and classification plays the first and foremost important role in capturing semantics in data and anchoring in translation as well as downstream study for history. However, NER in historical text has faced…

计算与语言 · 计算机科学 2023-06-27 Sojung Lucia Kim , Taehong Jang , Joonmo Ahn , Hyungil Lee , Jaehyuk Lee

The goal of this research was to find a way to extend the capabilities of computers through the processing of language in a more human way, and present applications which demonstrate the power of this method. This research presents a novel…

计算与语言 · 计算机科学 2013-01-17 Benjamin Englard

Political discourse datasets are important for gaining political insights, analyzing communication strategies or social science phenomena. Although numerous political discourse corpora exist, comprehensive, high-quality, annotated datasets…

‹ 上一页 1 2 3 10 下一页 ›