中文
相关论文

相关论文: Spanish Biomedical Crawled Corpus: A Large, Divers…

200 篇论文

In the recent years, transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack of corpora in…

This paper introduces the first version of the NUBes corpus (Negation and Uncertainty annotations in Biomedical texts in Spanish). The corpus is part of an on-going research and currently consists of 29,682 sentences obtained from…

计算与语言 · 计算机科学 2020-04-03 Salvador Lima , Naiara Perez , Montse Cuadros , German Rigau

We present a novel contribution to Spanish clinical natural language processing by introducing the largest publicly available clinical corpus, ClinText-SP, along with a state-of-the-art clinical encoder language model, RigoBERTa Clinical.…

计算与语言 · 计算机科学 2025-03-25 Guillem García Subies , Álvaro Barbero Jiménez , Paloma Martínez Fernández

We present a system that allows life-science researchers to search a linguistically annotated corpus of scientific texts using patterns over dependency graphs, as well as using patterns over token sequences and a powerful variant of boolean…

计算与语言 · 计算机科学 2020-06-09 Hillel Taub-Tabib , Micah Shlain , Shoval Sadde , Dan Lahav , Matan Eyal , Yaara Cohen , Yoav Goldberg

The BVS database (Health Virtual Library) is a centralized source of biomedical information for Latin America and Carib, created in 1998 and coordinated by BIREME (Biblioteca Regional de Medicina) in agreement with the Pan American Health…

计算与语言 · 计算机科学 2019-05-07 Felipe Soares , Martin Krallinger

This survey focuses in encoder Language Models for solving tasks in the clinical domain in the Spanish language. We review the contributions of 17 corpora focused mainly in clinical tasks, then list the most relevant Spanish Language Models…

计算与语言 · 计算机科学 2023-08-07 Guillem García Subies , Álvaro Barbero Jiménez , Paloma Martínez Fernández

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G…

计算与语言 · 计算机科学 2020-03-06 Liang Xu , Xuanwei Zhang , Qianqian Dong

Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet,…

We release a multilingual neural machine translation model, which can be used to translate text in the biomedical domain. The model can translate from 5 languages (French, German, Italian, Korean and Spanish) into English. It is trained…

计算与语言 · 计算机科学 2020-08-10 Alexandre Bérard , Zae Myung Kim , Vassilina Nikoulina , Eunjeong L. Park , Matthias Gallé

We present DepCC, the largest-to-date linguistically analyzed corpus in English including 365 million documents, composed of 252 billion tokens and 7.5 billion of named entity occurrences in 14.3 billion sentences from a web-scale crawl of…

计算与语言 · 计算机科学 2018-03-01 Alexander Panchenko , Eugen Ruppert , Stefano Faralli , Simone Paolo Ponzetto , Chris Biemann

Whilst there has been growing progress in Entity Linking (EL) for general language, existing datasets fail to address the complex nature of health terminology in layman's language. Meanwhile, there is a growing need for applications that…

计算与语言 · 计算机科学 2020-10-09 Marco Basaldella , Fangyu Liu , Ehsan Shareghi , Nigel Collier

Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which raises…

Multi-word expressions (MWEs) are a hot topic in research in natural language processing (NLP), including topics such as MWE detection, MWE decomposition, and research investigating the exploitation of MWEs in other NLP fields such as…

计算与语言 · 计算机科学 2020-05-22 Lifeng Han , Gareth J. F. Jones , Alan F. Smeaton

We present a new release of the Czech-English parallel corpus CzEng 2.0 consisting of over 2 billion words (2 "gigawords") in each language. The corpus contains document-level information and is filtered with several techniques to lower the…

计算与语言 · 计算机科学 2020-07-08 Tom Kocmi , Martin Popel , Ondrej Bojar

This paper presents a new annotated corpus of 513 anonymized radiology reports written in Spanish. Reports were manually annotated with entities, negation and uncertainty terms and relations. The corpus was conceived as an evaluation…

计算与语言 · 计算机科学 2017-11-01 Viviana Cotik , Darío Filippo , Roland Roller , Hans Uszkoreit , Feiyu Xu

In this technical report, we present Skywork-13B, a family of large language models (LLMs) trained on a corpus of over 3.2 trillion tokens drawn from both English and Chinese texts. This bilingual foundation model is the most extensively…

\textbf{Background:} Dutch medical corpora are scarce, limiting NLP development. \\ \textbf{Methods:} We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. \\…

计算与语言 · 计算机科学 2026-04-29 B. van Es

In machine translation (MT), health is a high-stakes domain characterised by widespread deployment and domain-specific vocabulary. However, there is a lack of MT evaluation datasets for low-resource languages in this domain. To address this…

计算与语言 · 计算机科学 2025-10-07 Raphaël Merx , Hanna Suominen , Trevor Cohn , Ekaterina Vylomova

Biomedical concept normalization links concept mentions in texts to a semantically equivalent concept in a biomedical knowledge base. This task is challenging as concepts can have different expressions in natural languages, e.g.…

计算与语言 · 计算机科学 2018-07-10 Roland Roller , Madeleine Kittner , Dirk Weissenborn , Ulf Leser
‹ 上一页 1 2 3 10 下一页 ›