English
Related papers

Related papers: A lexicon obtained and validated by a data-driven …

200 papers

Text mining is about looking for patterns in natural language text, and may be defined as the process of analyzing text to extract information from it for particular purposes. In previous work, we claimed that compression is a key…

Digital Libraries · Computer Science 2007-05-23 Stuart Yeates , David Bainbridge , Ian H. Witten

Automated terminology extraction refers to the task of extracting meaningful terms from domain-specific texts. This paper proposes a novel machine learning approach to terminology extraction, which combines features from traditional term…

Computation and Language · Computer Science 2025-02-25 Andraž Repar , Nada Lavrač , Senja Pollak

This paper presents a lexical disambiguation system, initially developed for English and now adapted to French. This system associates a word with its meaning in a given context using electronic dictionaries as semantically annotated…

Digital Libraries · Computer Science 2016-08-16 Caroline Brun , Bernard Jacquemin , Frédérique Segond

Environmental experts have developed the DPSIR (Driver, Pressure, State, Impact, Response) framework to systematically study and communicate key relationships between society and the environment. Using this framework requires experts to…

Human-Computer Interaction · Computer Science 2025-06-23 Sam Yu-Te Lee , Cheng-Wei Hung , Mei-Hua Yuan , Kwan-Liu Ma

Evaluating ecological time series is critical for benchmarking model performance in many important applications, including predicting greenhouse gas fluxes, capturing carbon-nitrogen dynamics, and monitoring hydrological cycles. Traditional…

Artificial Intelligence · Computer Science 2025-05-21 Qi Cheng , Licheng Liu , Qing Zhu , Runlong Yu , Zhenong Jin , Yiqun Xie , Xiaowei Jia

Creating labeled natural language training data is expensive and requires significant human effort. We mine input output examples from large corpora using a supervised mining function trained using a small seed set of only 100 examples. The…

Computation and Language · Computer Science 2022-05-10 Mandar Joshi , Terra Blevins , Mike Lewis , Daniel S. Weld , Luke Zettlemoyer

This research on data extraction methods applies recent advances in natural language processing to evidence synthesis based on medical texts. Texts of interest include abstracts of clinical trials in English and in multilingual contexts.…

Computation and Language · Computer Science 2020-01-31 Lena Schmidt , Julie Weeds , Julian P. T. Higgins

Important data are locked in ancient literature. It would be uneconomic to produce these data again and today or to extract them without the help of text mining technologies. Vespa is a text mining project whose aim is to extract data on…

Information Retrieval · Computer Science 2015-04-24 Nicolas Turenne , Mathieu Andro , Roselyne Corbière , Tien T. Phan

We consider automatically identifying the defined term within a mathematical definition from the text of an academic article. Inspired by the development of transformer-based natural language processing applications, we pose the problem as…

Artificial Intelligence · Computer Science 2023-11-22 Shufan Jiang , Pierre Senellart

Plant species identification in the wild is a difficult problem in part due to the high variability of the input data, but also because of complications induced by the long-tail effects of the datasets distribution. Inspired by the most…

Computer Vision and Pattern Recognition · Computer Science 2021-06-07 Matthew R. Keaton , Ram J. Zaveri , Meghana Kovur , Cole Henderson , Donald A. Adjeroh , Gianfranco Doretto

This paper proposes an approach to environmental accounting useful for studying the feasibility of socio-economic systems in relation to the external constraints posed by ecological compatibility. The approach is based on a multi-scale…

Populations and Evolution · Quantitative Biology 2017-04-27 Pedro L. Lomas , Mario Giampietro

Pest identification is a crucial aspect of pest control in agriculture. However, most farmers are not capable of accurately identifying pests in the field, and there is a limited number of structured data sources available for rapid…

Artificial Intelligence · Computer Science 2023-08-08 Ruoling Peng , Kang Liu , Po Yang , Zhipeng Yuan , Shunbao Li

Existing models which generate textual explanations enforce task relevance through a discriminative term loss function, but such mechanisms only weakly constrain mentioned object parts to actually be present in the image. In this paper, a…

Computer Vision and Pattern Recognition · Computer Science 2017-11-20 Lisa Anne Hendricks , Ronghang Hu , Trevor Darrell , Zeynep Akata

In this article is analyzed technology of automatic text abstracting and annotation. The role of annotation in automatic search and classification for different scientific articles is described. The algorithm of summarization of natural…

Computation and Language · Computer Science 2019-05-08 Nataliya Shakhovska , Taras Cherna

Large, open datasets can accelerate ecological research, particularly by enabling researchers to develop new insights by reusing datasets from multiple sources. However, to find the most suitable datasets to combine and integrate,…

Digital Libraries · Computer Science 2025-10-07 Zehao Lu , Thijs L van der Plas , Parinaz Rashidi , W Daniel Kissling , Ioannis N Athanasiadis

We analyze the state of the art of content-based retrieval in Earth observation image archives focusing on complete systems showing promise for operational implementation. The different paradigms at the basis of the main system families are…

Information Retrieval · Computer Science 2014-05-23 Marco Quartulli , Igor G. Olaizola

Managing the semantic quality of the categorization in large textual datasets, such as Wikipedia, presents significant challenges in terms of complexity and cost. In this paper, we propose leveraging transformer models to distill semantic…

Computation and Language · Computer Science 2024-04-26 Zineddine Bettouche , Anas Safi , Andreas Fischer

This paper presents a novel method for parsing and vectorizing semi-structured data to enhance the functionality of Retrieval-Augmented Generation (RAG) within Large Language Models (LLMs). We developed a comprehensive pipeline for…

Databases · Computer Science 2024-05-09 Hang Yang , Jing Guo , Jianchuan Qi , Jinliang Xie , Si Zhang , Siqi Yang , Nan Li , Ming Xu

A high-quality content analysis is essential for retrieval functionalities but the manual extraction of key phrases and classification is expensive. Natural language processing provides a framework to automatize the process. Here, a…

Computation and Language · Computer Science 2013-07-01 Ulf Schöneberg , Wolfram Sperber

This work presents an Argument Mining process that extracts argumentative entities from clinical texts and identifies their relationships using token classification and Natural Language Inference techniques. Compared to straightforward…

Computation and Language · Computer Science 2025-06-17 Maitane Urruela , Sergio Martín , Iker De la Iglesia , Ander Barrena