English
Related papers

Related papers: The Open Language Archives Community and Asian Lan…

200 papers

MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval across 18 different languages, which collectively encompass…

The need for raw large raw corpora has dramatically increased in recent years with the introduction of transfer learning and semi-supervised learning methods to Natural Language Processing. And while there have been some recent attempts to…

Computation and Language · Computer Science 2022-01-19 Julien Abadji , Pedro Ortiz Suarez , Laurent Romary , Benoît Sagot

Numerous institutions and organizations need not only to preserve the material and publications they produce, but also have as their task (although it would be desirable it was an obligation) to publish, disseminate and make publicly…

Digital Libraries · Computer Science 2017-05-24 J. Federico Medrano

In this article I outline the ideas behind the Open Archives Initiative metadata harvesting protocol (OAIMH), and attempt to clarify some common misconceptions. I then consider how the OAIMH protocol can be used to expose and harvest…

Digital Libraries · Computer Science 2010-08-11 Simeon Warner

We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from CommonCrawl and previously unused web crawls from the…

Natural Language Processing (NLP) for low-resource languages remains fundamentally constrained by the lack of textual corpora, standardized orthographies, and scalable annotation pipelines. While recent advances in large language models…

Computation and Language · Computer Science 2026-02-10 Bonaventure F. P. Dossou , Henri Aïdasso

OLAF (Open Life Science Analysis Framework) is an open-source platform that enables researchers to perform bioinformatics analyses using natural language. By combining large language models (LLMs) with a modular agent-pipe-router…

Quantitative Methods · Quantitative Biology 2025-04-14 Dylan Riffle , Nima Shirooni , Cody He , Manush Murali , Sovit Nayak , Rishikumar Gopalan , Diego Gonzalez Lopez

OpenAlex is an open bibliographic database that has been proposed as an alternative to commercial platforms in a context defined by the aim of transforming science evaluation systems into more transparent sources based on open data. This…

Digital Libraries · Computer Science 2025-12-19 Ángel Borrego , Cristóbal Urbano

In this work, we revisit linguistic acceptability in the context of large language models. We introduce CoLAC - Corpus of Linguistic Acceptability in Chinese, the first large-scale acceptability dataset for a non-Indo-European language. It…

Computation and Language · Computer Science 2023-09-29 Hai Hu , Ziyin Zhang , Weifang Huang , Jackie Yan-Ki Lai , Aini Li , Yina Patterson , Jiahui Huang , Peng Zhang , Chien-Jer Charles Lin , Rui Wang

Despite the existence of numerous Optical Character Recognition (OCR) tools, the lack of comprehensive open-source systems hampers the progress of document digitization in various low-resource languages, including Bengali. Low-resource…

Accurately aligning contextual representations in cross-lingual sentence embeddings is key for effective parallel data mining. A common strategy for achieving this alignment involves disentangling semantics and language in sentence…

Computation and Language · Computer Science 2025-09-03 Dayeon Ki , Cheonbok Park , Hyunjoong Kim

It is a well-known fact that current AI-based language technology -- language models, machine translation systems, multilingual dictionaries and corpora -- focuses on the world's 2-3% most widely spoken languages. Recent research efforts…

Computation and Language · Computer Science 2023-07-26 Gábor Bella , Paula Helm , Gertraud Koch , Fausto Giunchiglia

Clarivate's Web of Science (WoS) and Elsevier's Scopus have been for decades the main sources of bibliometric information. Although highly curated, these closed, proprietary databases are largely biased towards English-language…

The proliferation of open large language models (LLMs) is fostering a vibrant ecosystem of research and innovation in artificial intelligence (AI). However, the methods of collaboration used to develop open LLMs both before and after their…

Software Engineering · Computer Science 2025-10-01 Johan Linåker , Cailean Osborne , Jennifer Ding , Ben Burtenshaw

Faced with a considerable lack of resources in African languages to carry out work in Natural Language Processing (NLP), Natural Language Understanding (NLU) and artificial intelligence, the research teams of NTeALan association has set…

Computation and Language · Computer Science 2021-04-01 Elvis Mboning Tchiaze

The metaphor studies community has developed numerous valuable labelled corpora in various languages over the years. Many of these resources are not only unknown to the NLP community, but are also often not easily shared among the…

Computation and Language · Computer Science 2025-03-11 Joanne Boisson , Arif Mehmood , Jose Camacho-Collados

The automatic detection of offensive language is a pressing societal need. Many systems perform well on explicit offensive language but struggle to detect more complex, nuanced, or implicit cases of offensive and hateful language. OLEA is…

Computation and Language · Computer Science 2022-11-01 Marie Grace , Xajavion "Jay" Seabrum , Dananjay Srinivas , Alexis Palmer

By the end of the late 90's the Open Archives Initiative needed direction to insure its improvement and thus, created the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH) standard. The movement showed a rise in…

Digital Libraries · Computer Science 2017-08-30 Arnaud Gaudinat , Jonas Beausire , Megan Fuss , Elisa Banfi , Julien Gobeill , Patrick Ruch

The Open Archives Initiative (OAI) has recently created the Object Reuse and Exchange (ORE) project that defines Resource Maps (ReMs) for describing aggregations of web resources. These aggregations are susceptible to many of the same…

Digital Libraries · Computer Science 2009-01-30 Frank McCown , Michael L. Nelson , Herbert Van de Sompel

This paper describes the growth of Open Access (OA) repositories and journals as reported by monitoring initiatives such as ROAR (Registry of Open Access Repositories), Open DOAR (Open Directory of Open Access Repositories), DOAJ (Directory…

Digital Libraries · Computer Science 2013-01-24 A. N. Zainab