中文
相关论文

相关论文: A lexicon obtained and validated by a data-driven …

200 篇论文

Text mining is about looking for patterns in natural language text, and may be defined as the process of analyzing text to extract information from it for particular purposes. In previous work, we claimed that compression is a key…

数字图书馆 · 计算机科学 2007-05-23 Stuart Yeates , David Bainbridge , Ian H. Witten

Automated terminology extraction refers to the task of extracting meaningful terms from domain-specific texts. This paper proposes a novel machine learning approach to terminology extraction, which combines features from traditional term…

计算与语言 · 计算机科学 2025-02-25 Andraž Repar , Nada Lavrač , Senja Pollak

This paper presents a lexical disambiguation system, initially developed for English and now adapted to French. This system associates a word with its meaning in a given context using electronic dictionaries as semantically annotated…

数字图书馆 · 计算机科学 2016-08-16 Caroline Brun , Bernard Jacquemin , Frédérique Segond

Environmental experts have developed the DPSIR (Driver, Pressure, State, Impact, Response) framework to systematically study and communicate key relationships between society and the environment. Using this framework requires experts to…

人机交互 · 计算机科学 2025-06-23 Sam Yu-Te Lee , Cheng-Wei Hung , Mei-Hua Yuan , Kwan-Liu Ma

Evaluating ecological time series is critical for benchmarking model performance in many important applications, including predicting greenhouse gas fluxes, capturing carbon-nitrogen dynamics, and monitoring hydrological cycles. Traditional…

人工智能 · 计算机科学 2025-05-21 Qi Cheng , Licheng Liu , Qing Zhu , Runlong Yu , Zhenong Jin , Yiqun Xie , Xiaowei Jia

Creating labeled natural language training data is expensive and requires significant human effort. We mine input output examples from large corpora using a supervised mining function trained using a small seed set of only 100 examples. The…

计算与语言 · 计算机科学 2022-05-10 Mandar Joshi , Terra Blevins , Mike Lewis , Daniel S. Weld , Luke Zettlemoyer

This research on data extraction methods applies recent advances in natural language processing to evidence synthesis based on medical texts. Texts of interest include abstracts of clinical trials in English and in multilingual contexts.…

计算与语言 · 计算机科学 2020-01-31 Lena Schmidt , Julie Weeds , Julian P. T. Higgins

Important data are locked in ancient literature. It would be uneconomic to produce these data again and today or to extract them without the help of text mining technologies. Vespa is a text mining project whose aim is to extract data on…

信息检索 · 计算机科学 2015-04-24 Nicolas Turenne , Mathieu Andro , Roselyne Corbière , Tien T. Phan

We consider automatically identifying the defined term within a mathematical definition from the text of an academic article. Inspired by the development of transformer-based natural language processing applications, we pose the problem as…

人工智能 · 计算机科学 2023-11-22 Shufan Jiang , Pierre Senellart

Plant species identification in the wild is a difficult problem in part due to the high variability of the input data, but also because of complications induced by the long-tail effects of the datasets distribution. Inspired by the most…

计算机视觉与模式识别 · 计算机科学 2021-06-07 Matthew R. Keaton , Ram J. Zaveri , Meghana Kovur , Cole Henderson , Donald A. Adjeroh , Gianfranco Doretto

This paper proposes an approach to environmental accounting useful for studying the feasibility of socio-economic systems in relation to the external constraints posed by ecological compatibility. The approach is based on a multi-scale…

种群与进化 · 定量生物学 2017-04-27 Pedro L. Lomas , Mario Giampietro

Pest identification is a crucial aspect of pest control in agriculture. However, most farmers are not capable of accurately identifying pests in the field, and there is a limited number of structured data sources available for rapid…

人工智能 · 计算机科学 2023-08-08 Ruoling Peng , Kang Liu , Po Yang , Zhipeng Yuan , Shunbao Li

Existing models which generate textual explanations enforce task relevance through a discriminative term loss function, but such mechanisms only weakly constrain mentioned object parts to actually be present in the image. In this paper, a…

计算机视觉与模式识别 · 计算机科学 2017-11-20 Lisa Anne Hendricks , Ronghang Hu , Trevor Darrell , Zeynep Akata

In this article is analyzed technology of automatic text abstracting and annotation. The role of annotation in automatic search and classification for different scientific articles is described. The algorithm of summarization of natural…

计算与语言 · 计算机科学 2019-05-08 Nataliya Shakhovska , Taras Cherna

Large, open datasets can accelerate ecological research, particularly by enabling researchers to develop new insights by reusing datasets from multiple sources. However, to find the most suitable datasets to combine and integrate,…

数字图书馆 · 计算机科学 2025-10-07 Zehao Lu , Thijs L van der Plas , Parinaz Rashidi , W Daniel Kissling , Ioannis N Athanasiadis

We analyze the state of the art of content-based retrieval in Earth observation image archives focusing on complete systems showing promise for operational implementation. The different paradigms at the basis of the main system families are…

信息检索 · 计算机科学 2014-05-23 Marco Quartulli , Igor G. Olaizola

Managing the semantic quality of the categorization in large textual datasets, such as Wikipedia, presents significant challenges in terms of complexity and cost. In this paper, we propose leveraging transformer models to distill semantic…

计算与语言 · 计算机科学 2024-04-26 Zineddine Bettouche , Anas Safi , Andreas Fischer

This paper presents a novel method for parsing and vectorizing semi-structured data to enhance the functionality of Retrieval-Augmented Generation (RAG) within Large Language Models (LLMs). We developed a comprehensive pipeline for…

数据库 · 计算机科学 2024-05-09 Hang Yang , Jing Guo , Jianchuan Qi , Jinliang Xie , Si Zhang , Siqi Yang , Nan Li , Ming Xu

A high-quality content analysis is essential for retrieval functionalities but the manual extraction of key phrases and classification is expensive. Natural language processing provides a framework to automatize the process. Here, a…

计算与语言 · 计算机科学 2013-07-01 Ulf Schöneberg , Wolfram Sperber

This work presents an Argument Mining process that extracts argumentative entities from clinical texts and identifies their relationships using token classification and Natural Language Inference techniques. Compared to straightforward…

计算与语言 · 计算机科学 2025-06-17 Maitane Urruela , Sergio Martín , Iker De la Iglesia , Ander Barrena