中文
相关论文

相关论文: esCorpius: A Massive Spanish Crawling Corpus

200 篇论文

The Scielo database is an important source of scientific information in Latin America, containing articles from several research domains. A striking characteristic of Scielo is that many of its full-text contents are presented in more than…

计算与语言 · 计算机科学 2019-05-07 Felipe Soares , Viviane Pereira Moreira , Karin Becker

Pretrained language models are now ubiquitous in Natural Language Processing. Despite their success, most available models have either been trained on English data or on the concatenation of data in multiple languages. This makes practical…

The rapid development of multilingual large language models (LLMs) highlights the need for high-quality, diverse, and well-curated multilingual datasets. In this paper, we introduce DCAD-2000 (Data Cleaning as Anomaly Detection), a…

计算与语言 · 计算机科学 2025-10-27 Yingli Shen , Wen Lai , Shuo Wang , Xueren Zhang , Kangyang Luo , Alexander Fraser , Maosong Sun

We present a computationally-grounded word similarity dataset based on two well-known Natural Language Processing resources; text corpora and knowledge bases. This dataset aims to fulfil a gap in psycholinguistic research by providing a…

计算与语言 · 计算机科学 2023-04-21 J. Goikoetxea , M. Arantzeta , I. San Martin

This work introduces Salamandra, a suite of open-source decoder-only large language models available in three different sizes: 2, 7, and 40 billion parameters. The models were trained from scratch on highly multilingual data that comprises…

ROOTS is a 1.6TB multilingual text corpus developed for the training of BLOOM, currently the largest language model explicitly accompanied by commensurate data governance efforts. In continuation of these efforts, we present the ROOTS…

This paper presents WanJuan-CC, a safe and high-quality open-sourced English webtext dataset derived from Common Crawl data. The study addresses the challenges of constructing large-scale pre-training datasets for language models, which…

Language models (LMs) have introduced a major paradigm shift in Natural Language Processing (NLP) modeling where large pre-trained LMs became integral to most of the NLP tasks. The LMs are intelligent enough to find useful and relevant…

计算与语言 · 计算机科学 2023-05-09 Abbas Raza Ali , Muhammad Ajmal Siddiqui , Rema Algunaibet , Hasan Raza Ali

Information about pretraining corpora used to train the current best-performing language models is seldom discussed: commercial models rarely detail their data, and even open models are often released without accompanying training data or…

Fake news detection is a challenging task aiming to reduce human time and effort to check the truthfulness of news. Automated approaches to combat fake news, however, are limited by the lack of labeled benchmark datasets, especially in…

计算与语言 · 计算机科学 2021-03-02 Inna Vogel , Jeong-Eun Choi , Meghana Meghana

The utilization of clinical reports for various secondary purposes, including health research and treatment monitoring, is crucial for enhancing patient care. Natural Language Processing (NLP) tools have emerged as valuable assets for…

计算与语言 · 计算机科学 2023-06-14 Iker de la Iglesia , Aitziber Atutxa , Koldo Gojenola , Ander Barrena

Most previous work on the recently developed language-modeling approach to information retrieval focuses on document-specific characteristics, and therefore does not take into account the structure of the surrounding corpus. We propose a…

信息检索 · 计算机科学 2007-05-23 Oren Kurland , Lillian Lee

We present an evaluation of text simplification (TS) in Spanish for a production system, by means of two corpora focused in both complex-sentence and complex-word identification. We compare the most prevalent Spanish-specific readability…

计算与语言 · 计算机科学 2023-08-16 Adrian de Wynter , Anthony Hevia , Si-Qing Chen

We conducted a detailed analysis on the quality of web-mined corpora for two low-resource languages (making three language pairs, English-Sinhala, English-Tamil and Sinhala-Tamil). We ranked each corpus according to a similarity measure and…

计算与语言 · 计算机科学 2024-06-17 Surangika Ranathunga , Nisansa de Silva , Menan Velayuthan , Aloka Fernando , Charitha Rathnayake

The objective of the PANACEA ICT-2007.2.2 EU project is to build a platform that automates the stages involved in the acquisition, production, updating and maintenance of the large language resources required by, among others, MT systems.…

计算与语言 · 计算机科学 2013-03-11 Núria Bel , Vassilis Papavasiliou , Prokopis Prokopidis , Antonio Toral , Victoria Arranz

Training models for the automatic correction of machine-translated text usually relies on data consisting of (source, MT, human post- edit) triplets providing, for each source sentence, examples of translation errors with the corresponding…

计算与语言 · 计算机科学 2018-03-21 Matteo Negri , Marco Turchi , Rajen Chatterjee , Nicola Bertoldi

The multilingual nature of the world makes translation a crucial requirement today. Parallel dictionaries constructed by humans are a widely-available resource, but they are limited and do not provide enough coverage for good quality…

计算与语言 · 计算机科学 2015-12-08 Krzysztof Wołk , Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such data from…

计算与语言 · 计算机科学 2015-09-30 Krzysztof Wołk , Krzysztof Marasek

The lack of wide coverage datasets annotated with everyday metaphorical expressions for languages other than English is striking. This means that most research on supervised metaphor detection has been published only for that language. In…

计算与语言 · 计算机科学 2022-10-25 Elisa Sanchez-Bayona , Rodrigo Agerri

Comparable corpus is a set of topic aligned documents in multiple languages, which are not necessarily translations of each other. These documents are useful for multilingual natural language processing when there is no parallel text…

计算与语言 · 计算机科学 2025-08-05 Motaz Saad , David Langlois , Kamel Smaili